Dragos - Staff Engineering Support Engineer
Requirements
• Demonstrated experience troubleshooting complex distributed systems in production environments — comfortable reading logs, tracing component interactions, and isolating root cause under pressure. • Hands-on experience with cloud infrastructure operations (AWS, Azure, or GCP) and container orchestration platforms (Kubernetes), including day-to-day fleet maintenance and incident response. • Solid understanding of Linux systems administration and command-line tooling. • Experience supporting Elasticsearch, PostgreSQL, or similar database platforms in a production environment. • Experience developing and maintaining monitoring and alerting infrastructure (e.g., Datadog, Prometheus, Grafana) — not just consuming dashboards, but building and tuning them. • Strong written communication skills with a demonstrated ability to produce clear, accurate technical documentation (RCAs, runbooks, postmortems) for both internal and customer-facing audiences. • Experience operating effectively in a customer-facing or customer-adjacent engineering role, with the judgment to balance urgency and thoroughness on competing priorities under SLA pressure.
Responsibilities
• Lead cross-team incident triage for high-impact customer outages and escalations, coordinating Engineering, Product, and Customer Experience response, and contributing to preliminary root cause analysis and postmortem documentation. • Lead cross-team incident triage • Develop and maintain observability for cloud-hosted customer deployments — build and refine system monitors, dashboards, and alerting to ensure proactive detection and fast response to platform health events. • Develop and maintain observability • Serve as the senior escalation point for complex support cases in EMEA, working cases that involve multi-component interactions, deep platform internals, and unusual failure modes beyond standard troubleshooting guides. • Serve as the senior escalation point • Translate field experience into documentation — author comprehensive scenario-based troubleshooting guides, playbooks, and knowledge base articles that Tier 1/2 support staff can act on independently. • Translate field experience into documentation • Drive pattern recognition across the customer base — identify recurring failure modes, surface trends to product teams, and actively collaborate on permanent fixes rather than workarounds. • Drive pattern recognition across the customer base
Benefits
• Salary: £76,000 • Competitive Equity Package • Comprehensive Benefits Plan • #LI-NH1 #LI-REMOTE
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT