wagey.ggwagey.gg
31,341  jobs31,341  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,341)/Site Reliability Engineer Role(176)/Epic Kids Inc. (7) - Senior Site Reliability Engineer
Epic Kids Inc.

Epic Kids Inc. - Senior Site Reliability Engineer

Remote - USA1mo ago
RemoteSeniorNACloud ComputingSite Reliability EngineerBashPythonMandarinGCPDockerKubernetesHelmJenkinsTerraformNew RelicGoogle GKEAirflowDagsterCFPSAFe

Requirements

• Bachelor's degree or higher in Computer Science, Software Engineering, or a related field. • 5+ years of experience in infrastructure, platform, DevOps, or a related engineering role, with a track record of measurably improving production reliability—including defining SLOs, reducing incident frequency or MTTR, and eliminating recurring failure modes. • Hands-on experience with Google Cloud Platform (GCP), including GCE, GCS, VPC, IAM, Cloud Monitoring, and related services. • Experience with Docker and Kubernetes (GKE), including containerizing workloads, Helm, and cluster fundamentals. • Experience with CI/CD pipelines such as GitHub Actions, ArgoCD, Jenkins, or similar tools. • Experience with an observability platform such as New Relic, including metrics, logging, alerting, and dashboards. • Proficiency with Terraform for managing infrastructure as code. • Scripting or programming experience with Python, Bash, or similar languages. • Experience operating workflow orchestration platforms such as Dagster or Airflow as a service for data or platform teams. • Familiarity with PromRelay for metrics forwarding and alert routing. • Familiarity with the operational footprint of data platforms, including warehouse infrastructure, job schedulers, and batch workloads. • Experience working within distributed or global engineering teams. • Working knowledge of compliance frameworks such as SOC 2, FERPA, and COPPA, as well as GRC tools. • Proficiency in Mandarin Chinese is a plus.

Responsibilities

• Drive the reliability of Epic's infrastructure—set and track SLOs/SLIs, reduce toil, and engineer out recurring instability. • Build and operate the cloud infrastructure and container platform for high availability, scalability, and cost efficiency—including workload scheduling, autoscaling, networking, and graceful failure handling. • Maintain and improve CI/CD pipelines for fast, safe delivery across engineering teams. • Own and evolve the observability stack—metrics, logs, traces, dashboards, and alerts. • Manage infrastructure as code across the organization, with a focus on consistency, change safety, and reproducibility. • Own platform security practices—including secrets management, IAM policies, and network segmentation. • Support compliance-aware infrastructure practices—including vulnerability management, access reviews, audit-evidence flows, and incident-response readiness. • Participate in a frequent on-call rotation; drive incident response, blameless post-mortems, and follow-through on systemic fixes. • Partner with product and data engineering teams to troubleshoot platform issues and guide developers on infrastructure best practices.

Benefits

• Work alongside talented teammates in a collaborative, supportive, and global environment. • Enjoy the flexibility of a fully remote, U.S.-based position. • Help build and scale the infrastructure powering millions of young readers around the world.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

BranchBranch - Senior Site Reliability Engineer (SRE)1mo ago
·Remote - USA·$364k+/year
RemoteNASeniorCloud ComputingSite Reliability EngineerJavaSpring BootTerraformDockerGoKubernetesBashPythonGoogle Pub/SubRedisPrometheusMySQLGrafanaGCP
Stack AVStack AV - Senior Site Reliability Engineer2mo ago
·Remote - Pittsburgh, PA or Remote
RemoteNASeniorCloud ComputingGovernmentSite Reliability EngineerBashPythonLinuxGCPAWS
AmwellAmwell - Senior Site Reliability Engineer2mo ago
·US - Remote - Hybrid·$129k - $140k/year + Equity
In OfficeNASeniorMental HealthCloud ComputingSite Reliability EngineerBashPythonPuppetTerraformAnsible
PinterestPinterest - Site Reliability Engineer II, tvScientific2mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform
Bedrock Ocean ExplorationBedrock Ocean Exploration - Senior Site Reliability Engineer, Robotics & Cloud Infrastructure2mo ago
·Brooklyn, New York, USA·$164k - $220k/year + Equity
In OfficeNASeniorCloud ComputingRoboticsSite Reliability EngineerTeam ManagementCustomer OnboardingBashPythonGo
synthesiasynthesia - Senior Site Reliability Enigneer2mo ago
·Remote - USA
RemoteNASeniorCloud ComputingSite Reliability EngineerTeam ManagementAWSKubernetesMongoDBPython
AuthZedAuthZed - Sr. Site Reliability Engineer3mo ago
·Remote - USA·$150k - $195k/year + Equity
RemoteNASeniorInsuranceCloud ComputingSite Reliability EngineerJavaRubyPythonGoDocker
GitLabGitLab - Senior Site Reliability Engineer, Environment Automation6mo ago
·Remote - Canada·$124k - $266k/year + Equity
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformKubernetesAnsibleRubyGo
talkiatrytalkiatry - Senior Site Reliability Engineer1mo ago
·Remote - Americas
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformPythonTypeScriptTeam ManagementPrometheusDatadogAWSGrafanaNode.jsReactKubernetesDocumentation

Browse more by category

Show 176 moreSite Reliability EngineerShow 358 moreBashShow 5,059 morePythonShow 262 moreMandarinShow 1,263 moreGCPShow 892 moreDockerShow 1,670 moreKubernetesShow 134 moreHelmShow 155 moreJenkinsShow 920 moreTerraform
Privacy·Terms··Contact·FAQ·Wagey on X