encord - Senior SRE Engineer
Requirements
• Experience on hands-on SRE, DevOps, platform engineering experience or similar in a production environment. • Strong fundamentals in designing, building and maintaining resilient distributed and/or high performance systems • Solid understanding of networking, operating systems and database technologies • Experience with observability fundamentals — metrics, logs, traces, and alerting. • Comfortable with on-call rotations and incident management. • We are technology agnostic at Encord and not looking for experience across all of these — as long as you're open to learning, please apply. • Backend: Python and Rust • Frontend: TypeScript and React • Deployment: Kubernetes • Infrastructure: GCP
Responsibilities
• Performance & Capacity — Profile and optimise services handling large-scale data pipelines; perform capacity planning for storage and compute-intensive workloads. Work with squads to establish performance benchmarks and expectations • Collaboration — Partner closely with backend and ML engineers to improve deployment pipelines (CI/CD), review infrastructure changes, and champion reliability best practices. • Reliability & Availability — Define and own SLIs/SLOs/SLAs for critical services; build alerting, runbooks, and incident response processes; lead postmortems with a blameless culture. • Infrastructure & Cloud — Design, deploy, and maintain cloud infrastructure on GCP and AWS; manage Kubernetes clusters, networking, and storage at petabyte scale. • Automation & Tooling — Work to improve developer productivity and guide and review automation and tooling efforts across the engineering group. • Observability — Instrument services with distributed tracing, logging, and metrics (Prometheus, Grafana, OpenTelemetry, Datadog or similar); build infrastructure, define best practices and work with each squad to ensure every service is observable before it goes to production.
Benefits
• Competitive salary, commission, and meaningful equity in a high-growth startup • Strong in-person culture — most of the team works from our London office 4+ days/week • 25 days annual leave + UK public holidays • Annual learning & development budget • Travel for customer visits, events, and conferences across the UK and Europe • Company lunches twice a week • Monthly socials & bi-annual team offsites
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT