wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/Infrastructure Engineer Role(198)/radiant (15) - Infrastructure Tooling & Observability Engineer( UK)
radiant

radiant - Infrastructure Tooling & Observability Engineer( UK)

London, England, United Kingdom1mo ago
In OfficeSeniorEMEAInfrastructure EngineerQA EngineerRubyGoRailsPerformance ManagementAnsible

Requirements

• Degree in Computer Science/Software Engineering, or equivalent experience • 6–8 years of experience in infrastructure engineering, DevOps, SRE, and/or software engineering roles, with a strong focus on operational systems. • Proven experience in at least one recent DevOps or software engineering role, building or maintaining production infrastructure tooling or platform systems. • Experience working in large-scale or distributed infrastructure environments (hyperscale, enterprise, or similarly complex systems). • Strong programming ability in at least one of: Ruby (Rails), Go, or similar systems languages, with willingness and ability to work across multiple languages and codebases. • Hands-on experience with infrastructure automation tools such as Ansible and orchestration platforms such as AWX. • Strong experience with observability systems, including the Grafana stack (Prometheus, Loki, Mimir, and Grafana Alloy). • Familiarity with low-level telemetry and infrastructure protocols such as SNMP and syslog. • Experience working with Kubernetes or similar orchestration platforms in production environments. • Understanding of API design and integration patterns, particularly REST-based services and service-to-service communication. • Experience building and maintaining CI/CD pipelines, including GitHub Actions and self-hosted runners. • Strong understanding of operational reliability concepts, including monitoring, alerting, capacity planning, and incident response. • Comfortable working closely with SRE, Platform Engineering, and infrastructure teams to translate operational needs into maintainable software systems. • Kubernetes Certified Administrator • Cloud-native observability training courses, attendance at industry conferences in this field • LPI/LPIC certification

Responsibilities

• Design, build, and evolve internal tooling and observability platforms that support large-scale infrastructure operations across distributed environments. • Develop systems that turn high-volume telemetry (logs, metrics, events) into actionable insight, improving visibility, alerting quality, and operational decision-making. • Translate SRE reliability requirements into scalable, production-ready software solutions, including automation for incident detection, prevention, and remediation. • Drive automation across infrastructure operations, reducing manual effort in areas such as environment provisioning, cluster onboarding, inventory management, and lifecycle workflows. • Build tooling for capacity management, performance testing, benchmarking, and automated collection and analysis of results. • Contribute to Continual Service Improvement (CSI) initiatives by identifying operational inefficiencies and delivering durable engineering solutions. • Work closely with SRE and infrastructure engineering teams to embed observability and reliability into core platform workflows. • Interface with Platform Engineering teams to ensure tooling aligns with broader orchestration and infrastructure strategy. • Integrate and extend existing systems written in Ruby/Rails and Go, contributing to a consistent and maintainable engineering ecosystem. • Develop and maintain automation workflows using Ansible and AWX. • Support CI/CD-driven operational tooling, including GitHub Actions and self-hosted runners.

Benefits

• This role goes beyond traditional monitoring. You will build the internal control plane that transforms high-volume telemetry—logs, metrics, and events—into actionable insight for engineering and operations teams. Your work will improve observability across infrastructure systems, strengthen signal quality, and help teams understand and respond to system behaviour in real time. • Working closely with SRE and infrastructure engineering teams, you will translate reliability goals into scalable, production-grade tooling. This includes frameworks for observability, alerting, anomaly detection, capacity planning, and service health tracking. • A key focus of the role is automation. You will help eliminate manual processes across infrastructure operations, including environment provisioning, cluster onboarding, inventory management, and recurring operational workflows. You will also contribute to performance engineering initiatives, building tooling for testing, benchmarking, and automated results collection at scale. • You will play a central role in turning SRE reliability initiatives into reusable engineering solutions, including automated remediation systems and tooling that reduces operational toil while improving system resilience. • You can also expect: • Exposure to large-scale distributed infrastructure systems • Opportunities to shape foundational internal platforms • A collaborative, engineering-led culture with strong ownership • High-impact work spanning observability, automation, and reliability • Close partnership with SRE and infrastructure engineering teams • A fast-moving environment where tooling directly improves operational performance

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

olixolix - Infrastructure Engineer1mo ago
·London, England, United Kingdom - Hybrid·£88k/year/year + Equity
In OfficeEMEASeniorCloud ComputingDeveloper ToolsInfrastructure EngineerTerraformBashAnsiblePythonKubernetes
Orcrist TechnologiesOrcrist Technologies - (GPU) Infrastructure Engineer1mo ago
·Remote - Germany
RemoteEMEASeniorArtificial IntelligenceMaterialsInfrastructure EngineerTerraformAnsibleCUDAKubernetesPulumi
ChainalysisChainalysis - Senior Infrastructure Engineer2mo ago
·Tel Aviv, Israel
In OfficeEMEASeniorCryptocurrencyCloud ComputingInfrastructure EngineerRustPythonGoKubernetesAWS
writerwriter - Infrastructure engineer (UK)2mo ago
·London, England, United Kingdom - Hybrid·Equity
In OfficeEMEASeniorCloud ComputingArtificial IntelligenceInfrastructure EngineerJavaGoPythonTerraformAWS
TelnyxTelnyx - Senior Infrastructure Engineer (Core)5mo ago
·Remote - EU (Dublin/Amsterdam office or Remote), LATAM
RemoteEMEASeniorSoftwareInfrastructure EngineerPythonAnsiblePuppetTerraformKubernetes
NamespaceNamespace - Infrastructure Software Engineer2mo ago
·Zurich, Switzerland
In OfficeEMEASoftware EngineerInfrastructure EngineerKubernetesTerraformAnsibleGo
FractileFractile - Silicon Infrastructure Engineer1mo ago
·London - Hybrid·Equity
In OfficeEMEAArtificial IntelligenceDeveloper ToolsSemiconductorsInfrastructure EngineerPythonDockerKubernetesTerraformAnsible
Sporty GroupSporty Group - Lead Hybrid QA Engineer (Europe only)2mo ago
·Remote - Global
RemoteEMEAStaffDeveloper ToolsQA EngineerReportingTeam ManagementTeam LeadershipPerformance ManagementFull Stack
TelnyxTelnyx - Infrastructure Engineer (Core)2mo ago
·Remote - Dublin, Ireland; Ho Chi Minh City, Vietnam; Bangalore, India; Warsaw, Poland; Amsterdam, Netherlands
RemoteEMEAMidSoftwareInfrastructure EngineerPythonAnsiblePuppetTerraformKubernetes

Browse more by category

Show 198 moreInfrastructure EngineerShow 198 moreQA EngineerShow 366 moreRubyShow 1,778 moreGoShow 146 moreRailsShow 1,172 morePerformance ManagementShow 160 moreAnsible
Privacy·Terms··Contact·FAQ·Wagey on X