Valtech - Senior Site Reliability Engineer
Requirements
• You are someone with 5 years of experience in the field of software engineering, DevOps engineering, QA engineering and/or cloud engineering, of which at least the last 2 years as a dedicated Site Reliability Engineer. You feel comfortable taking the lead, making decisions, and know how to mobilize and motivate people to set things in motion. In your current role, people come to you for advice on what to look for to determine the robustness of their production environments, advice for reliable deployment procedures, assistance in analysis of failure scenarios, and ideas on how to mitigate or remediate those. • To be considered for this role, you must meet the following essential qualifications: • You are assertive with good communicative skills, capable of taking the lead and coaching a development team to make the right choices. • You have experience with incident management in a production environment of a public-facing online service with high business value and preferably high traffic in a 24x7 fashion. • You have experience in working in corporate environments. • You have experience programming and scripting. • You have at least basic knowledge of serverless services in one or more public cloud providers (AWS, Azure, GCP). • You have extensive knowledge of and experience with various monitoring systems, amongst which APM systems such as Datadog, New Relic, Dynatrace, Prometheus, and Grafana. • You have knowledge of and experience with various pipelining tools, such as GitHub, Azure DevOps, GitLab, Jenkins. • You have knowledge of and experience with microservices-related technology: Docker, Kubernetes. • You have a good conceptual understanding of software architecture and system thinking. • You have worked as an engineer in a DevOps context. • You have an excellent command of English (C1 or above). • Datadog (or APM equivalent). • Java / Springboot. • Kubernetes / EKS. • Have worked within the context of publicly accessible, highly available eCommerce platforms. • Have experience working in an international context with on- and off-shore teams. • Commitment to reaching all kinds of people • We design experiences that work for all kinds of people - and that starts with our own teams. At Valtech, we’re intentional about building an inclusive culture where everyone feels supported to grow, thrive and achieve their goals. No matter your background, you belong here. Explore our Diversity & Inclusion site to see how we’re creating a more equitable Valtech for all.
Responsibilities
• Work with teams to define SLIs and SLOs. • Create systems for observability. • Work with teams to analyze failure scenarios and possible mitigations. • (Assisting to) create runbooks to remediate or prevent failure scenarios. • Reduce work that does not add value. • Participate and facilitate incident management, including on-call duty.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT