wagey.ggwagey.gg
30,538  jobs30,538  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,538)/Site Reliability Engineer Role(169)/CSC Generation (10) - Site Reliability Engineer
CSC Generation

CSC Generation - Site Reliability Engineer

Remote - Costa Rica1w ago
RemoteMidLATAMCloud ComputingE-commerceSite Reliability EngineerHR ManagerBashPythonTypeScriptNode.jsKubernetesTerraformTeam ManagementAnsibleAzureCDKClaudeLinuxPrometheusGrafanaLokiKustomizeHelmEcommerceAWSGCP

Requirements

• 3+ years of experience supporting containerized production services, preferably running Kubernetes • 3+ years of experience with Infrastructure as Code (Terraform, AWS CDK, Ansible, etc.) • 3+ years of cloud experience operating in Google Cloud Platform and/or AWS (multi-cloud stack; Azure/Entra exposure is a plus) • Comfortable diagnosing issues and shipping bug fixes directly to application code (not just infrastructure) to keep services reliable and stable • Comfortable performing deep dives across both infrastructure and application/software git repositories to trace issues end-to-end • Proficient with AI-assisted coding tools (e.g., Claude Code, GitHub Copilot) and MCP-based agents, used to accelerate investigation, code review, and automation • Strong knowledge of scripting and programming languages (Bash, Python, and TypeScript/Node.js) • Experience managing Linux (any major distribution) in production environments • Excellent understanding of internet application protocols (DHCP, DNS, HTTPS, SSH, etc.) • Understanding of how DevOps (CI/CD) and SRE practices (SLOs, SLIs) apply to daily work • Hands-on experience with observability tooling (Grafana, Prometheus, Loki, OpenSearch, or equivalents) and SLI/SLO instrumentation • Experience with GitOps and Kubernetes packaging (ArgoCD, Helm, Kustomize) • Proactively track emerging technology trends and developments, evaluating which ones are worth bringing into engineering practice • Bachelor's degree in computer science or similar, or equivalent experience • Advanced-level English communication skills, both verbal and written • Experience using AI coding assistants (Claude Code, Codex, GitHub Copilot) to build fixes, write automation, and ship application and infrastructure code improvements • Familiarity with PCI-scoped or other regulated environments • Previous experience working in ecommerce environments • Professional certifications: GCP, CKA, or AWS • Hiring Manager Interview - Conversation with the Site Reliability Manager focused on your SRE experience, approach to incident management, and team fit. • Technical/Case Discussion - A deeper dive into infrastructure, observability, and problem-solving scenarios relevant to the role. • Reference Checks - Conducted in parallel with the final stages where possible. • Offer - We move quickly for the right candidate. • Interview process is subject to change. Any updates will be communicated promptly and clearly.

Responsibilities

• Work on service resiliency, performance tuning, and system design across Backcountry's platform • Drive resolution of critical incidents and ensure fixes are methodically implemented through postmortems • Leverage AI-assisted engineering tools (Claude Code, GitHub Copilot, MCP-based agents) to investigate, automate, and ship fixes across infrastructure and application repositories • Reduce toil by designing and implementing automation • Partner with other Site Reliability Engineers, developers, and architects to evaluate and implement best practices for current and future workloads • Monitor system health and capacity, taking proactive action to fix problems before they occur • Collaborate with engineering teams to build, deploy, and support features • Build and maintain observability (metrics, logs, traces, profiles) and SLI/SLO instrumentation for Backcountry services • Participate in FinOps initiatives across GCP and AWS, including capacity planning and committed-use discount strategy • Participate in the on-call support rotation within the SRE team

Benefits

• The people who do best here are builders. They take ownership, move fast, and want to see the direct impact of their work. • Cross-Functional Impact: Your work directly affects platform reliability for every customer and every team that depends on Backcountry's systems. • Modern Tech Stack: Work across a multi-cloud environment (GCP and AWS) with modern observability tooling, AI-assisted engineering, and GitOps workflows. • End-to-End Ownership: Own projects from investigation through implementation — you will ship automation, improve resiliency, and see the results in production. • Competitive Benefits: We offer an attractive benefits package including primarily remote work, private medical and life insurance, additional paid time off, monthly allowances and reimbursements, employee discounts, and opportunities for professional growth.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

PinterestPinterest - Site Reliability Engineer II, tvScientific1mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform
EnumerateEnumerate - Senior Site Reliability Engineer1mo ago
·Remote - LATAM·$48k - $48k/year
RemoteLATAMSeniorCloud ComputingSoftwareSite Reliability EngineerGovernanceDocumentationBashPythonKubernetes
KrakenKraken - Sr. Site Reliability Engineer5mo ago
·Remote - LATAM·$104k - $104k/year
RemoteLATAMSeniorCryptocurrencyCloud ComputingSite Reliability EngineerRustAWSPythonDockerKubernetes
talkiatrytalkiatry - Senior Site Reliability Engineer1w ago
·Remote - Americas
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformPythonTypeScriptTeam ManagementPrometheusDatadogAWSGrafanaNode.jsReactKubernetesDocumentation
SecurityScorecardSecurityScorecard - Senior Site Reliability Engineer1w ago
·Remote - Brazil·Equity
RemoteLATAMSeniorArtificial IntelligenceSite Reliability EngineerBashGoPythonKubernetesJenkinsTerraformHelmPrometheusGoogle GKEGovernancePulumiGrafanaDatadogKafkaMLOps
AccelaAccela - Site Reliability Engineer 21mo ago
·Remote - Based - US·$125k - $145k/year + Equity
RemoteNAMidInsuranceCloud ComputingSite Reliability EngineerBashPythonChange ManagementAzureKubernetes
Okta, Inc.Okta, Inc. - Staff Site Reliability Engineer1mo ago
·Bengaluru, India
In OfficeAPACStaffCloud ComputingSoftwareSite Reliability EngineerKubernetesTimeline ManagementAWSGCPHelmTerraformPythonGoGoogle GKELinuxAnsible
ValtechValtech - Senior Site Reliability Engineer1w ago
·Remote - North Macedonia
RemoteWWSeniorCloud ComputingE-commerceSite Reliability EngineerAWSGCPAzurePrometheusDynatraceDatadogNew RelicJenkinsGrafanaDockerKubernetesJavaCoaching

Browse more by category

Show 169 moreSite Reliability EngineerShow 124 moreHR ManagerShow 327 moreBashShow 4,739 morePythonShow 1,943 moreTypeScriptShow 705 moreNode.jsShow 1,520 moreKubernetesShow 829 moreTerraformShow 2,713 moreTeam ManagementShow 146 moreAnsible
Privacy·Terms··Contact·FAQ·Wagey on X