• Write high-quality, maintainable software — primarily in Python, but we value engineering ability over language familiarity.
• high-quality, maintainable software
• Have a strong background in scalable infrastructure, including:
• scalable infrastructure
• Containerization and orchestration (e.g. Docker, Kubernetes)
• Infrastructure-as-code and deployment (e.g. Terraform, CI/CD pipelines)
• Monitoring and logging frameworks (e.g. Datadog, Prometheus, OpenTelemetry)
• Understand and implement ML Ops best practices, including:
• ML Ops best practices
• Model versioning and rollback strategies
• Automated evaluation and drift detection
• Scalable model and agent serving infrastructure (e.g. vLLM, Triton, BentoML)
• Deploy and maintain LLM and agentic workflows in production, including:
• LLM and agentic workflows
• Monitoring cost, latency, and performance
• Capturing traces for analysis and debugging
• Optimizing prompt/response flows with real-time data access
• Demonstrate strong ownership and pragmatism, balancing infrastructure elegance with iterative delivery and measurable impact.
• strong ownership and pragmatism