glia - Senior Software Engineer, SRE / Observability Tooling
Requirements
• Infrastructure: AWS, Kubernetes (AWS EKS), Istio, EFK • Persistence: Amazon Aurora Serverless for Postgres, RabbitMQ, Amazon RDS • Cache: Amazon ElastiCache • Monitoring & Observability: DataDog with a focus on dashboards and alerts for system health • CI/CD: Github Actions, ArgoCD, Jenkins, Helm, with a focus on automation and pipeline optimization. • Infrastructure as Code: Terraform • Additionally, our Engineering teams use: • Backend: Python, Elixir, Node.js http://node.js, Ruby, Go • Frontend: Javascript and React.js • Native mobile SDKs: Java and Swift • Expert-level proficiency with AWS and Kubernetes (EKS), particularly in areas of observability, networking, and auto-scaling. • Experience with modern observability platforms (e.g., DataDog, Prometheus) and a deep understanding of metrics, logging, and tracing. • Deep, practical understanding of Site Reliability Engineering (SRE) principles (SLOs, error budgets, toil reduction). • Demonstrable experience analyzing and troubleshooting large-scale distributed systems. • Strong software development skills in a language like Python or Go, used to build operational tools, services, or automation. • Expertise in designing and operating robust CI/CD pipelines for a microservices architecture (e.g., using ArgoCD, Github Actions, Helm). • A systematic, data-driven approach to problem-solving and root cause analysis. • Proficiency in using AI tools thoughtfully, maintaining ownership of the final output while recognizing the tools' limitations. • Glia is an equal-opportunity employer. Glia does not discriminate against any employee or applicant because of race, creed, color, religion, gender, sexual orientation, gender identity/expression, national origin, disability, age, genetic information, veteran status, marital status, pregnancy or related condition (including breastfeeding), or any other basis protected by law. • The Glia Talent Acquisition team uses @glia.com http://glia.com and @ [email protected] http://gliatalent.com email addresses for coordinating interviews, providing updates, and sending documents. • Our hiring process involves an introduction, practical and team interviews, and a decision and offer. For more information, visit our Recruitment Privacy Notice page https://www.glia.com/eu-recruitment-privacy-notice or contact our talent team via [email protected]
Responsibilities
• Developing standards, infrastructure and automation for dashboards, alerts, and monitors as code. • Partnering with development teams to establish production readiness and operational readiness. • Building the tooling and templates teams use to define, measure, and report on Service Level Objectives (SLOs) and Service Level Indicators (SLIs) for their services. • Developing tooling to automate observability and operational workflows, eliminating manual toil for engineering teams. • Building and improving the incident response tooling and workflows that help teams resolve outages faster and learn from them. • Our Collaboration Model • Glia Engineering is remote-first, spanning Canada, Portugal, Poland, and Estonia. The Observability team is based in Estonia, with optional offices in Tallinn and Tartu. We thrive on flexible remote collaboration, but we still bring the whole team together in Estonia twice a year for in-person innovation and connection.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT