• Implementing the improvements to the reliability, fault tolerance, scalability, and performance of our infrastructure
• Managing incidents using your technical know-how to involve the appropriate teams and automate away manual practices
• Providing support to our critical services by responding to automated alerts through our on-call rotation
• Define and maintain SLIs, SLOs,SLA, and error budgets to guide reliability decisions
• Improve observability across our systems (metrics, logs, tracing) to reduce time to detection and resolution
• Make production issues easier to detect, troubleshoot, and resolve
• Improving monitoring, alerting, dashboards, tracing and runbooks for critical services
• Leading postmortems and follow-up actions to reduce repeat incidents