Gremlin - Data Scientist, AI/ML
Requirements
• Experience with chaos engineering, site reliability engineering, or distributed systems • Background in agentic AI systems or large-scale causal inference in production • Experience standing up MLOps tooling such as model serving, monitoring, or feature store infrastructure • Working in Remote first environments • Has been on-call and participated in an incident management program • The role does not offer sponsorship employment benefits. • *If you don’t think you meet all of the criteria above but still are interested in the job, please apply. Nobody checks every box, we’re looking for candidates that are particularly strong in a few areas, and have some interest and capabilities in others.
Responsibilities
• Analyze Gremlin’s proprietary dataset of millions of chaos engineering experiments to identify failure patterns, root causes, and resilience signals across complex distributed systems • Pretraining and fine-tuning machine learning models that automatically detect, classify, and explain failures observed during chaos experiments • Build intelligent systems that deliver automated remediation recommendations, and eventually orchestration, by learning from historical experiment outcomes and system behavior • Develop scalable data pipelines and feature stores to process, enrich, and serve large volumes of experiment data for both model training and real-time inference • Collaborate closely with platform engineers and SREs to integrate AI-driven failure analysis and remediation capabilities directly into Gremlin’s core product • Translate insights from millions of chaos experiments into AI-powered features that help customers automatically understand blast radius, pinpoint root causes, and accelerate recovery • Research and productionize novel ML approaches, including causal AI and agentic systems, that turn raw chaos experiment data into automated, reliable remediation strategies • We’ll expect you to have: • Experience as a self-driven and collaborative problem solver with strong communication skills • 5+ years professional experience building and productionizing machine learning, ideally for distributed systems, infrastructure, or DevOps and SRE use cases with more overall years of experience in software development. • Hands-on experience with techniques such as causal inference, graph ML, time-series modeling, or reinforcement learning • Experience building data pipelines and feature stores that support both offline training and real-time inference • Experience with agile development environments and practices • Strong advocate and practitioner of rigorous experimentation, model evaluation, and engineering best practices • Comfort partnering with platform engineers and SREs to turn research into shipped product features • Strong at breaking down ambiguous problems into concrete actions and milestones
Benefits
• We expect the salary range for this role to be $220,000 - $290,000. We recognize that salary varies from person to person depending on level of experience and we welcome direct conversations about it. The final offer will vary based on assessment of a candidate's skills and ability and our budget and market data. • Gremlin offers competitive total compensation packages including 401k Matching, Equity and other benefits such as flexible time off and paid company holidays. • Gremlin is a team of industry veterans and people eager to learn from one another. We set the standard for reliability and equip leading organizations with the mindset and expertise needed to drive reliability improvements that move the world forward. We’re backed by top-tier investors Index Ventures, Amplify Partners, and Redpoint Ventures. Our customers love us, and we’re thrilled to be a partner in their success.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT