deepgram - Senior Software Engineer - Model Evaluation & AI Systems
Requirements
• BS, MS, or PhD in Computer Science, AI, Applied Math, or a related field, or equivalent experience. • 5+ years of professional software or QA engineering experience, with a track record of shipping test infrastructure or evaluation systems (senior candidates with significantly deeper experience welcome). • Solid backend/scripting experience in a language such as Python, Rust, Go, or similar. • Experience designing and building automated test pipelines, evaluation frameworks, or data-processing systems. • Strong analytical skills and comfort reasoning about metrics, thresholds, and statistical variation in results — able to distinguish real regressions from noise. • Ability to take charge of ambiguous technical challenges and communicate effectively across research, engineering, and product teams. • Hands-on experience evaluating modern AI systems such as LLMs, RAG pipelines, agents, or multimodal models, including model behavior analysis. • Experience with React Native or other cross-platform mobile frameworks for building tooling that's accessible beyond the desktop. • Experience building or improving evaluation frameworks, benchmarks, or ML infrastructure used by other teams or external users. • A strong appreciation for evaluation quality — correctness, reproducibility, and consistency across environments. • Experience with voice, audio, speech recognition, or real-time systems, and familiarity with metrics like WER, MOS, or latency/TTFB. • Prior involvement in open-source projects, through contributions, reviews, maintenance, or community engagement. • Experience acting as a technical bridge across teams or platforms (evaluation, training, inference, agent frameworks), combining architectural understanding with clear communication and influence. • Familiarity with cloud infrastructure, containerized/ephemeral environments, and monitoring tooling (e.g. Grafana, canaries, anomaly detection).
Responsibilities
• Define and build evaluation methodologies for Deepgram's models, spanning speech-to-text, text-to-speech, and emerging LLM, RAG, agent, and multimodal systems. • Design, build, and maintain automated evaluation pipelines across batch and streaming (e.g. WER, runaway/hallucination detection, latency and time-to-first-byte), with a focus on correctness, reproducibility, and ease of adoption. • Build scalable, reproducible evaluation infrastructure — harnesses, orchestration, and result-aggregation pipelines — running against production models and, where needed, large GPU clusters. • Translate Research benchmarks and expected model metrics into automated, enforceable pass/fail gates. • Build and operate canaries and continuous-monitoring systems that detect quality regressions in production before they reach customers. • Partner with DevOps/Infra to stand up ephemeral test environments and results-aggregation infrastructure. • Work alongside Research, model training, inference, and product teams to provide trusted evaluation signals that inform release and optimization decisions. • Integrate evaluation and quality gates into CI/CD so quality is verified continuously, not manually. • Help raise the bar through code reviews, technical design discussions, and strong engineering and QA practices.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT