wagey.ggwagey.gg
30,244  jobs30,244  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,244)/Site Reliability Engineer Role(165)/Airalo (21) - Senior Site Reliability Engineer
Airalo

Airalo - Senior Site Reliability Engineer

Spain / United Kingdom+ Equity1mo ago
RemoteSeniorEMEACloud ComputingTelecommunicationsSite Reliability EngineerGoJavaPythonPrometheusDatadog

Requirements

• Deep understanding of observability principles and tools such as: Prometheus, Datadog, OpenTelemetry and similar. • Experience with leading incident management and complex postmortem analysis. • Experience and interest in managing infrastructure as code (Terraform). • Experience with chaos engineering and other techniques for testing system resilience. • Experience with CI/CD tools such as GitHub Actions for automated delivery. • Proficiency in at least one programming language (Python, Go, Java, etc.) for building automation and internal tooling. • Event-driven architecture experience (SNS, SQS etc) • Ability to work independently and collaboratively in a fast-paced environment. • Team player and open to new ideas. • Good communication skills and fluency in English. • Prior experience with Scrum and other agile methods. • Certification in relevant areas such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), or similar. • Prior experience with Telco Core Networks (e.g., 5G/LTE Packet Core, IMS, Signaling) and low-latency networking. • Experience with AI-driven SRE tools for anomaly detection and improvements • Contributions to open-source SRE projects or communities. • Prior work experience in telecommunications. • Deep understanding of eSIM and GSMA related technologies and services. • If you are interested in this position, please apply via the link. • By applying, you acknowledge and agree that, in case of successful application, Airalo may request to run background checks as a condition for entering into an agreement with you. Rest assured that these checks will only occur upon your prior consent and at the end of the selection process, and will be strictly limited to what is allowed under the laws that are applicable to you. All data that you share or that we collect in connection with such checks will be processed in accordance with our Privacy Policy, available here: www.airalo.com/more-info/privacy-policy?srsltid=AfmBOooBT0rXAj1FaNelZ3VfN0wvhwzvAoxdtHnOKSVETpiSjiXVuycy • We sincerely thank all applicants in advance for submitting their interest in this opportunity. Airalo is an equal-opportunity employer and values diversity, equity & inclusion. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We are committed to providing reasonable accommodations upon request for individuals with disabilities throughout our job interview process.

Responsibilities

• Lead the design of scalable, fault-tolerant and self-healing systems in a multi-region AWS environment. • Define and track Service Level Objectives (SLOs) and Service Level Indicators (SLIs) to drive architectural decisions and error budget policies. • Conduct blameless post-incident reviews to uncover systemic root causes and implement long-term preventive measures. • Identify patterns of manual work and lead the development of internal tools/automation to permanently eliminate them. • Develop and maintain automated runbooks and playbooks for common operational tasks and complex incident response. • Shift from simple monitoring to deep observability, ensuring high cardinality data leads to proactive actionable insights. • Proactively identify and mitigate operational risks through chaos engineering and architecture reviews. • Work with software engineers to design systems for reliability, scalability, and maintainability from the early stages of the SDLC. • Continuously evaluate and optimize system performance, capacity, and cost efficiency. • Beyond just participating, you will refine the on-call experience to reduce alert fatigue, improve MTTR, and ensure sustainable rotation health. • Must-haves: • Bachelor’s degree in Computer Engineering or a similar discipline. • 5+ years of experience as a Site Reliability Engineer or in a similar role. • 3+ years of experience with AWS services including strong knowledge of container orchestration. • 2+ years of Kubernetes experience • Deep understanding of observability principles and tools such as: Prometheus, Datadog, OpenTelemetry and similar. • Experience with leading incident management and complex postmortem analysis. • Experience and interest in managing infrastructure as code (Terraform). • Experience with chaos engineering and other techniques for testing system resilience. • Experience with CI/CD tools such as GitHub Actions for automated delivery. • Proficiency in at least one programming language (Python, Go, Java, etc.) for building automation and internal tooling. • Event-driven architecture experience (SNS, SQS etc) • Ability to work independently and collaboratively in a fast-paced environment. • Team player and open to new ideas. • Good communication skills and fluency in English. • Good to have: • Prior experience with Scrum and other agile methods. • Certification in relevant areas such as AWS Certified DevOps Engineer, Certified Kubernetes Administrator (CKA), or similar. • Prior experience with Telco Core Networks (e.g., 5G/LTE Packet Core, IMS, Signaling) and low-latency networking. • Experience with AI-driven SRE tools for anomaly detection and improvements • Contributions to open-source SRE projects or communities. • Prior work experience in telecommunications. • Deep understanding of eSIM and GSMA related technologies and services. • If you are interested in this position, please apply via the link. • By applying, you acknowledge and agree that, in case of successful application, Airalo may request to run background checks as a condition for entering into an agreement with you. Rest assured that these checks will only occur upon your prior consent and at the end of the selection process, and will be strictly limited to what is allowed under the laws that are applicable to you. All data that you share or that we collect in connection with such checks will be processed in accordance with our Privacy Policy, available here: www.airalo.com/more-info/privacy-policy?srsltid=AfmBOooBT0rXAj1FaNelZ3VfN0wvhwzvAoxdtHnOKSVETpiSjiXVuycy • We sincerely thank all applicants in advance for submitting their interest in this opportunity. Airalo is an equal-opportunity employer and values diversity, equity & inclusion. We do not discriminate on the basis of race, religion, national origin, gender, sexual orientation, age, marital status, veteran status, or disability status. We are committed to providing reasonable accommodations upon request for individuals with disabilities throughout our job interview process.

Benefits

• Hi, I'm Daniele, VP of Engineering at Airalo! • Engineering drives Airalo’s eSIM platform. We build the product that lets millions of people connect instantly across the globe. The challenges are exciting: high-scale systems, carrier integrations, and products spanning both B2C and B2B. What matters most to us is creating an environment where engineers can do their best work. Real ownership, autonomy, and a direct link between what you ship and business outcomes are key. You’ll work with smart, motivated people who take their craft seriously. If you want to build things that matter at global scale, this is where you do it. • Airalo’s fully remote Engineering team in Spain is growing. In this role, you'll tackle complex technical challenges across our product ecosystem, helping build, innovate, and scale the platform that keeps millions of travellers connected worldwide. • We are looking for a Senior Site Reliability Engineer to join our growing engineering team. • We are a company that values SRE principles and practices. We believe in empowering our SREs to make data-driven decisions, automate operational tasks, and continuously improve the reliability of our systems. We foster a blameless culture where everyone is encouraged to learn from mistakes and share knowledge. If you are passionate about building and maintaining highly reliable systems, we would love to hear from you! • Participating in our on-call rotation is a core expectation of this role. It's essential for maintaining 24/7 service reliability across our global operations, ensuring our systems remain resilient and our customers experience uninterrupted service, regardless of time zone or geography. • Paid Rotation: We offer standby fees + overtime pay. • Delayed Start: No on-call duties for your first 6 months. • Rest & Recovery: Guaranteed rest periods and flexible hours following night incidents. • Shared Load: Rotations are split (Weekdays vs. Weekends) to minimize fatigue. • Please refer to the On-Call Policy in the Airalo Handbook for full details: airalo-public.notion.site/our-approach-to-engineering-on-call-policy

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

AxonAxon - Sr Site Reliability Engineer I6mo ago
·London, England, United Kingdom
In OfficeEMEASeniorCloud ComputingSoftwareSite Reliability EngineerGoC#JavaPythonTerraform
NICENICE - Site Reliability Engineer3mo ago
·Remote - United Kingdom
RemoteEMEACloud ComputingSite Reliability EngineerGoSplunkDatadogBashPython
replitreplit - Senior Site Reliability Engineer2mo ago
·Remote - Europe
RemoteEMEASeniorCloud ComputingSite Reliability EngineerGoPythonReportingKubernetesGCP
ValtechValtech - Senior Site Reliability Engineer1w ago
·Remote - Poland
RemoteEMEASeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching
ValtechValtech - Senior Site Reliability Engineer1w ago
·Portugal - Remote - Hybrid
In OfficeEMEASeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching
AxonAxon - Site Reliability Engineer1w ago
·London, England, United Kingdom·Equity
In OfficeEMEAMidCloud ComputingSoftwareSite Reliability EngineerAWSAzureKubernetesPythonJavaC#GoDocumentation
AuthZedAuthZed - Sr. Site Reliability Engineer1mo ago
·Remote - USA·$150k - $195k/year + Equity
RemoteNASeniorInsuranceCloud ComputingSite Reliability EngineerJavaRubyPythonGoDocker
playonsportsplayonsports - PlayOn - Senior Site Reliability Engineer2mo ago
·Remote - USA *·Equity
RemoteNASeniorCloud ComputingSoftwareSite Reliability EngineerJavaC++GoChange ManagementPython
DittoDitto - Senior Site Reliability Engineer3mo ago
·Remote - USA·$156k - $288k/year + Equity
RemoteNASeniorCloud ComputingSoftwareSite Reliability EngineerGoRustC++JavaPython

Browse more by category

Show 165 moreSite Reliability EngineerShow 1,653 moreGoShow 1,437 moreJavaShow 4,673 morePythonShow 184 morePrometheusShow 166 moreDatadog
Privacy·Terms··Contact·FAQ·Wagey on X