wagey.ggwagey.gg
30,535  jobs30,535  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,535)/Site Reliability Engineer Role(169)/Sporty Group (19) - Site Reliability Engineer
Sporty Group

Sporty Group - Site Reliability Engineer

EMEA1w ago
In OfficeMidEMEACloud ComputingSite Reliability EngineerDBABashRustPythonAWSKubernetesHelmTerraformLokiPrometheusGrafanaMemcachedRedisLinuxNode.jsJavaScriptJavaSpring BootMongoDBMySQLPostgreSQLValkeyKafkaDockerVectorJenkinsCloudflareMentoring

Requirements

• 3+ years DevOps / SRE / platform engineering experience • Must be based in Europe • Experience independently leading the planning and deployment of a project • Experienced with cloud platforms, especially AWS, including solid knowledge of how to utilise cloud resources to fulfil the demand from other teams and production • Strong understanding of Kubernetes and container orchestration, with experience in EKS and GitOps tooling such as ArgoCD and Helm being highly valued • Experience with Infrastructure-as-Code, particularly Terraform • Proficiency in scripting and automation with Bash, Python, or Golang; experience with Rust is a plus • Hands-on experience with observability stacks covering metrics, logs, distributed traces, and profiling, for example Prometheus, Loki, Tempo, Pyroscope, and OpenTelemetry • Experience with real user monitoring (RUM), with familiarity in Grafana Faro or OpenTelemetry SDK instrumentation being a plus • Proven on-call and incident response experience, comfortable triaging production issues under pressure, leading post-mortems, and driving follow-up actions • Ability to design and maintain alert frameworks that minimise noise, prevent alert fatigue, and avoid waterfall alerting patterns • Experience defining SLIs and SLOs and using them to inform reliability work • Familiarity with service mesh concepts is a plus, as we are actively evaluating Cilium-based service mesh in non-production environments • Solid networking knowledge, especially the TCP / IP stack and HTTP protocol • Experience handling high HTTP request volumes and designing systems for high availability and high traffic environments • A strong understanding of cache, including CDN, HTTP cache, Redis / Memcached • Excellent troubleshooting skills, including Linux OS issue diagnosis and OS parameter optimisation, JVM optimisation would be highly advantageous • Our stack • Languages: Java / Spring Boot, Node.js, Python, JavaScript • Database: Aurora MySQL & PostgreSQL, MongoDB, MySQL Community • Cache: ElastiCache, Redis, Valkey • Messaging: Apache RocketMQ, AutoMQ, Kafka • Networking & Proxy: Nginx, Kong, Cilium, eBPF • Orchestration & GitOps: Docker, Kubernetes (EKS), ArgoCD, Helm • Computing & Storage: AWS EC2, VPC, AWS Lambda, EBS, S3 • CI/CD: Jenkins, GitHub Actions • Metrics: Prometheus, Mimir, Grafana, Alertmanager • Logs: Loki, Vector • Traces: Tempo, OpenTelemetry, Alloy • Profiling: Pyroscope • RUM: Grafana Faro, OpenTelemetry SDK • Infrastructure as Code: Terraform • CDN & Edge: Cloudflare, AWS CloudFront

Responsibilities

• Work with a team of DevOps and DBA professionals • Improve existing infrastructure and processes across the countries we’re deployed in, as well as streamlining processes to deploy to new countries in the future • Continuously improve Kubernetes platform stability and efficiency, with a focus on optimising resource utilisation, reducing costs, and streamlining environment provisioning through GitOps-first practices • Monitor and maintain cloud infrastructure through autoscaling, alerting pipelines, and Grafana dashboards covering metrics, logs, traces, and real user monitoring (RUM) • Own weekend on-call operations, triaging and responding to production incidents, performing root cause analysis, and driving post-incident reviews • Design and manage alert pipelines to ensure actionable signal quality, with attention to preventing alert fatigue, waterfall alerting, and notification flooding • Define and maintain SLIs and SLOs for critical services, and use them to drive reliability improvements and on-call prioritisation • Take ownership and responsibility for our cloud operation activities • Liaise with external security agencies for annual audits as well as perform our own internal security sweeps • Aid in reconfiguring existing architecture to allow for rapid deployments to new countries • Mentoring less experienced team members

Benefits

• Sporty is a remote first company in pursuit of sustainability • A competitive salary + individual performance based bonuses every quarter • 28 days paid annual leave • Our core working hours are 10am-3pm in your local time zone with flexibility outside of this • Referral bonuses & flash bonuses • Top of the line equipment • Annual company retreats to provide great internal networking opportunities

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

AlpacaAlpaca - Site Reliability Engineer2mo ago
·Remote - EMEA·Equity
RemoteEMEAMidFintechCloud ComputingSite Reliability EngineerDBAGoLinuxPythonKubernetesPostgreSQL
ClearScore Technology LimitedClearScore Technology Limited - Site Reliability Engineer1w ago
·London, England, United Kingdom
In OfficeEMEACloud ComputingSite Reliability EngineerPythonLinuxAWSKubernetesDocker SwarmNomadGoJenkinsSpinnakerElasticsearchLokiPrometheusDatadogGrafanaKafkaTerraform
AxonAxon - Site Reliability Engineer1w ago
·London, England, United Kingdom·Equity
In OfficeEMEAMidCloud ComputingSoftwareSite Reliability EngineerAWSAzureKubernetesPythonJavaC#GoDocumentation
airappsairapps - Site Reliability Engineer (SRE)3mo ago
·London, London Metropolitain Area, UK·€55k - €68k/year
In OfficeEMEAMidCloud ComputingSite Reliability EngineerBashGoTerraformPulumiPython
ValtechValtech - Senior Site Reliability Engineer1w ago
·Remote - Poland
RemoteEMEASeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching
ValtechValtech - Senior Site Reliability Engineer1w ago
·Portugal - Remote - Hybrid
In OfficeEMEASeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching
Sporty GroupSporty Group - DevOps Engineer1mo ago
·Remote - Global - Europe *
RemoteEMEAMidCloud ComputingDevOps EngineerDBAAWSKubernetesLokiPrometheusGrafana
PinterestPinterest - Site Reliability Engineer II, tvScientific1mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform
RemoteRemote - Senior Site Reliability Engineer1w ago
·Remote - EMEA·$53k - $53k/year
RemoteEMEASeniorCloud ComputingTransportationSite Reliability EngineerTeam LeadKubernetesDockerAWSReportingTerraformPrometheusGrafanaBashLinuxElixirPythonObservableNode.jsBack-end

Browse more by category

Show 169 moreSite Reliability EngineerShow 43 moreDBAShow 327 moreBashShow 616 moreRustShow 4,738 morePythonShow 2,772 moreAWSShow 1,520 moreKubernetesShow 119 moreHelmShow 829 moreTerraformShow 26 moreLoki
Privacy·Terms··Contact·FAQ·Wagey on X