Epic Kids Inc. - Senior Site Reliability Engineer
Requirements
• Bachelor's degree or higher in Computer Science, Software Engineering, or a related field. • 5+ years of experience in infrastructure, platform, DevOps, or a related engineering role, with a track record of measurably improving production reliability—including defining SLOs, reducing incident frequency or MTTR, and eliminating recurring failure modes. • Hands-on experience with Google Cloud Platform (GCP), including GCE, GCS, VPC, IAM, Cloud Monitoring, and related services. • Experience with Docker and Kubernetes (GKE), including containerizing workloads, Helm, and cluster fundamentals. • Experience with CI/CD pipelines such as GitHub Actions, ArgoCD, Jenkins, or similar tools. • Experience with an observability platform such as New Relic, including metrics, logging, alerting, and dashboards. • Proficiency with Terraform for managing infrastructure as code. • Scripting or programming experience with Python, Bash, or similar languages. • Experience operating workflow orchestration platforms such as Dagster or Airflow as a service for data or platform teams. • Familiarity with PromRelay for metrics forwarding and alert routing. • Familiarity with the operational footprint of data platforms, including warehouse infrastructure, job schedulers, and batch workloads. • Experience working within distributed or global engineering teams. • Working knowledge of compliance frameworks such as SOC 2, FERPA, and COPPA, as well as GRC tools. • Proficiency in Mandarin Chinese is a plus.
Responsibilities
• Drive the reliability of Epic's infrastructure—set and track SLOs/SLIs, reduce toil, and engineer out recurring instability. • Build and operate the cloud infrastructure and container platform for high availability, scalability, and cost efficiency—including workload scheduling, autoscaling, networking, and graceful failure handling. • Maintain and improve CI/CD pipelines for fast, safe delivery across engineering teams. • Own and evolve the observability stack—metrics, logs, traces, dashboards, and alerts. • Manage infrastructure as code across the organization, with a focus on consistency, change safety, and reproducibility. • Own platform security practices—including secrets management, IAM policies, and network segmentation. • Support compliance-aware infrastructure practices—including vulnerability management, access reviews, audit-evidence flows, and incident-response readiness. • Participate in a frequent on-call rotation; drive incident response, blameless post-mortems, and follow-through on systemic fixes. • Partner with product and data engineering teams to troubleshoot platform issues and guide developers on infrastructure best practices.
Benefits
• Work alongside talented teammates in a collaborative, supportive, and global environment. • Enjoy the flexibility of a fully remote, U.S.-based position. • Help build and scale the infrastructure powering millions of young readers around the world.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT