wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/Site Reliability Engineer Role(176)/Develocity (3) - Staff Site Reliability Engineer
Pro members applied to this job 36 hours before you saw itGet Pro ›
Develocity

Develocity - Staff Site Reliability Engineer

Remote - Europe (GMT)+ Equity2d ago
RemoteStaffEMEACloud ComputingSoftwareSite Reliability EngineerBashPythonCoachingNew Hire OnboardingAWSKubernetesDocumentationPrometheusTerraformGrafanaJavaKotlin

Requirements

• We're building a new SRE team and looking for founding members to help shape how we operate. As a Lead SRE, you’ll be a technical and operational leader for reliability across Develocity. You’ll help define our SRE vision, set standards for how we operate production services, and mentor other SREs as the team grows. This is a hands-on role with broad influence across engineering, cloud platform, and customer-facing teams. • The SRE team will be responsible for the reliability, performance, and availability of Develocity instances serving paying customers, open-source projects, and public-facing services, plus supporting infrastructure like artifact registries. • You'll work on our internally-built Cloud Application Platform, Kubernetes on AWS, and develop deep expertise in it. When incidents happen, you'll troubleshoot issues across the stack, from application to infrastructure. You'll collaborate with the Cloud Platform team to improve the tooling you depend on, and with engineering teams to build reliability into how we ship software. If you like automating things and hate doing the same task twice, you'll fit in well. • You'll be part of a distributed, remote-first team that values asynchronous communication and written documentation. Strong self-direction and clear communication across time zones are essential. • 7+ years in SRE, DevOps, or an equivalent role operating production services at scale. • Experience leading reliability initiatives across multiple teams or services. • Demonstrated ability to influence technical direction without direct authority. • Experience designing and operating systems with SLOs and error budgets, and exercising strong judgment in balancing reliability, velocity, and cost. • Strong Kubernetes experience in production environments. • Cloud infrastructure expertise, preferably AWS (EKS, RDS, S3, EC2). • Proficiency with observability tools (Prometheus, Grafana) and Infrastructure as Code (Terraform). • Track record of incident management and response in a 24/7 on-call environment. • Scripting proficiency (Python, Bash) for automation. • Strong written and verbal English communication skills. • Experience as a founding or early SRE establishing practices in a growing SaaS organization. • Familiarity with Develocity. • JVM language experience (Java, Kotlin). • Experience with customer-facing and executive-level incident communications.

Responsibilities

• Operate and maintain all Develocity instances and supporting services in production. • Define and evolve SRE standards, practices, and operating models, including on-call, incident response, postmortems, and SLOs. • Participate in a follow-the-sun on-call rotation, acting as a technical escalation point for complex or high-severity incidents. • Lead incident response and blameless retrospectives, ensuring learnings result in measurable reliability improvements. • Set reliability priorities using risk, customer impact, business goals, SLOs, and error budgets. • Identify systemic reliability risks and continuously evolve Develocity’s SaaS operations as the platform and customer base grow. • Lead and influence architectural and design reviews to ensure reliability, scalability, and operability. • Drive automation across deployment, upgrades, monitoring, self-healing, recovery, and operational workflows. • Build and maintain comprehensive observability for all managed services, including logging, metrics, tracing, and alerting. • Own disaster recovery, backups, and business continuity planning and execution. • Partner with engineering leadership to balance feature delivery with reliability and operational excellence. • Mentor and coach SREs, supporting technical growth and strong operational practices. • Help onboard new SREs and contribute to hiring by defining and assessing SRE excellence at Develocity. • Communicate clearly with customers during incidents and maintenance windows. • Optimize performance, resource utilization, and operational costs.

Benefits

• A ground-floor role in a new SRE team - you'll shape how we do things, not inherit someone else's decisions. • Real ownership of production systems used by engineers at companies you've heard of. • Direct interaction with customers when things go wrong (and when they go right). • A culture that values automation over heroics. • In-person meetings, such as our annual company offsite and team meetings. • Work from home in a remote-first environment. • Competitive salaries and equity grants. • Location • Location • Remote from anywhere in Europe (GMT). • While our team works remotely and is spread across the globe, we deeply value daily interactions and collaboration.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

DevelocityDevelocity - Senior Site Reliability Engineer2d ago
·Remote - Europe (GMT)·Equity
RemoteEMEASeniorCloud ComputingSoftwareSite Reliability EngineerBashPythonAWSKubernetesPrometheusDocumentationTerraformGrafanaJavaKotlin
AxonAxon - Site Reliability Engineer1w ago
·London, England, United Kingdom·Equity
In OfficeEMEAMidCloud ComputingSoftwareSite Reliability EngineerAWSAzureKubernetesPythonJavaC#GoDocumentation
AxonAxon - Sr Site Reliability Engineer I6mo ago
·London, England, United Kingdom
In OfficeEMEASeniorCloud ComputingSoftwareSite Reliability EngineerGoC#JavaPythonTerraform
ValtechValtech - Senior Site Reliability Engineer1w ago
·Remote - Poland
RemoteEMEASeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching
ValtechValtech - Senior Site Reliability Engineer1w ago
·Portugal - Remote - Hybrid
In OfficeEMEASeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching
ClearScore Technology LimitedClearScore Technology Limited - Site Reliability Engineer1w ago
·London, England, United Kingdom
In OfficeEMEACloud ComputingSite Reliability EngineerPythonLinuxAWSKubernetesDocker SwarmNomadGoJenkinsSpinnakerElasticsearchLokiPrometheusDatadogGrafanaKafkaTerraform
Ping IdentityPing Identity - Staff Site Reliability Engineer1w ago
·Remote - UK
RemoteEMEAStaffCloud ComputingSoftwareSite Reliability EngineerGCPReportingKubernetesGoDockerGit
RedditReddit - Staff Site Reliability Engineer - Site Experience2mo ago
·Remote - UK
RemoteEMEAStaffCloud ComputingSite Reliability EngineerGoPythonPerformance ManagementLinuxKubernetes
RedditReddit - Staff Site Reliability Engineer2mo ago
·Dublin, Ireland
In OfficeEMEAStaffCloud ComputingSite Reliability EngineerGoPythonPerformance ManagementLinuxKubernetes

Browse more by category

Show 176 moreSite Reliability EngineerShow 358 moreBashShow 5,063 morePythonShow 2,689 moreCoachingShow 111 moreNew Hire OnboardingShow 3,038 moreAWSShow 1,674 moreKubernetesShow 5,261 moreDocumentationShow 233 morePrometheusShow 921 moreTerraform
Privacy·Terms··Contact·FAQ·Wagey on X