• Degree in engineering, computer science, or a quantitative/physical science — or equivalent practical experience.
• 5+ years of software/ML engineering, including 2+ years building LLM-based systems that run in production.
• You have designed evaluations for LLM/agent systems — eval sets, quality metrics, human-expert or LLM-judge pipelines — and can walk us through one (e.g., promptfoo, Braintrust, LangSmith, DeepEval, or your own harness).
• You have instrumented, monitored, and debugged live AI services (e.g., OpenTelemetry, Arize Phoenix, Langfuse, Datadog, or similar).
• Strong Python; able to independently build and deploy services.
• ## Strong pluses (not required — you'll have room and support to pick these up on the job)
• Search/RAG: vector databases, keyword search, knowledge graphs, reranking, hybrid retrieval
• MCP (Model Context Protocol) or agent-tool ecosystem experience
• Materials science, chemistry, or patent/IP domain exposure
• Structured information extraction from technical documents (tables, compositions, specs)