wagey.ggwagey.gg
30,816  jobs30,816  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,816)/Site Reliability Engineer Role(171)/onebrief (19) - Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided
Pro members applied to this job 36 hours before you saw itGet Pro ›
onebrief

onebrief - Senior Site Reliability Engineer (Arlington, VA) - Secret Clearance Required - Relocation Provided

Remote - United States2d ago
RemoteSeniorNACloud ComputingSite Reliability EngineerRecruiterELKAWSTerraformDatadogGrafanaKubernetesGoAnsibleLinkerdIstioTypeScriptLokiPrometheusJenkinsBashPythonFront-end

Requirements

• Observability: Grafana stack, ELK, or Datadog • Infrastructure as Code (Terraform, Ansible) and cloud experience (AWS or AWS GovCloud) • Kubernetes cluster design and operations • Designing meaningful SLIs/SLOs with error budgets for distributed systems • GitOps practices and toolchains • DoD environments and compliance frameworks (RMF, STIGs, ICD 503) • Service mesh (Istio, Linkerd) • On-prem virtualization (VMware, Proxmox, Nutanix, Hyper-V) • Relevant certs (AWS DevOps Engineer, CKA/CKAD) • Notice to Third Party Recruitment Agencies • Please note that Onebrief does not accept unsolicited resumes from recruiters or employment agencies. In the absence of an executed Recruitment Services Agreement, there will be no obligation to any referral compensation or recruiter fee. In the event a recruiter or agency submits a resume or candidate without an agreement Onebrief explicitly reserves the right to pursue and hire those candidate(s) without any financial obligation to the recruiter or agency. Any unsolicited resumes, including those submitted to hiring managers, shall be deemed the property of Onebrief.

Responsibilities

• You'll help make our production application reliable, scalable, and secure by improving the software itself, not just the systems it runs on. Day to day that looks like: • Improving the application: Work directly in the codebase (primarily TypeScript) to fix reliability and performance problems at the source. You'll partner with product engineers on design decisions, review code with reliability and security in mind, and treat "make the app better" as a first-class part of the job rather than something you hand off. • Building observability that developers actually use: Design and run our monitoring, logging, and alerting (Prometheus, Loki, Alloy, Grafana). The goal is alerts and dashboards tied to real application behavior, so teams catch issues before users do. • Owning reliability targets: Define and measure SLIs and SLOs, wire up alerting that feeds them, and be the person who can say what "reliable" means for our systems and prove it with data. • Leading incident response: Act as incident responder, and incident commander when needed. Run blameless post-mortems (AARs) that find the actual root cause and turn it into a code or process fix so it doesn't happen again. • Automating away toil: Spot the repetitive operational work and write software to kill it. Share what works with other teams, including those running in air-gapped environments, and help them get production-ready. • WHAT WE LOOK FOR • An active Secret clearance • 5+ years in software engineering, SRE, or a related role, with real time spent writing and shipping application code • Strong TypeScript (or comparable modern language experience with willingness to work primarily in TypeScript) • Solid grasp of the full SDLC: design, code review, testing, release, and how reliability fits into each stage • Experience with incident response, root cause analysis, and turning findings into lasting fixes • A collaborator who works well across product, platform, and DevOps teams and shares context openly • TECHNICAL EXPERTISE • Application development in TypeScript (Node and/or a modern front-end framework) • CI/CD: building and maintaining pipelines (GitHub Actions, GitLab CI/CD, Jenkins) • Testing and quality practices as part of the delivery process • Comfort with at least one of Python, Go, or Bash for tooling and automation • Working knowledge of containers and Kubernetes (enough to debug and deploy, not necessarily to stand up clusters from scratch) • Networking fundamentals and secure configuration basics

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

onebriefonebrief - Senior Site Reliability Engineer (Arlington, VA) - Relocation Provided3w ago
·Remote - United States·$180k - $220k/year
RemoteNASeniorCloud ComputingSite Reliability EngineerRecruiterGoBashPythonLinkerdAWS
talkiatrytalkiatry - Senior Site Reliability Engineer1w ago
·Remote - Americas
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformPythonTypeScriptTeam ManagementPrometheusDatadogAWSGrafanaNode.jsReactKubernetesDocumentation
Flock SafetyFlock Safety - Site Reliability Engineer III, Aviation5mo ago
·United States·$140k - $165k/year + Equity
In OfficeNAMidCloud ComputingInternet of ThingsSite Reliability EngineerRecruiterDashboard CreationDocumentationTerraformPrometheusGrafana
Veeam SoftwareVeeam Software - Senior Site Reliability Engineer- FedRamp1w ago
·Remote - USA·$173k - $173k/year
RemoteNASeniorCloud ComputingArtificial IntelligenceSite Reliability EngineerGoJavaC#TypeScriptAWSAzurePrometheusGrafanaELKTerraformKubernetesDocumentationGitB2BPulumiCloseCosmosElastic Stack
AmwellAmwell - Senior Site Reliability Engineer1mo ago
·US - Remote - Hybrid·$129k - $140k/year + Equity
In OfficeNASeniorMental HealthCloud ComputingSite Reliability EngineerBashPythonPuppetTerraformAnsible
GitLabGitLab - Senior Site Reliability Engineer, Environment Automation4mo ago
·Remote - Canada·$124k - $266k/year + Equity
RemoteNASeniorCloud ComputingSite Reliability EngineerTerraformKubernetesAnsibleRubyGo
PinterestPinterest - Site Reliability Engineer II, tvScientific1mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform
Honeycomb.ioHoneycomb.io - Pipeline - Senior Site Reliability Engineer1mo ago
·Remote - UK·£128k/year/year + Equity
RemoteEMEASeniorBankingCloud ComputingSite Reliability EngineerRecruiterBaseAWSKubernetesTerraformHelm
Honeycomb.ioHoneycomb.io - Senior Site Reliability Engineer1mo ago
·Remote - Ireland·€140.59/hour/year + Equity
RemoteEMEASeniorBankingCloud ComputingSite Reliability EngineerRecruiterBaseAWSKubernetesTerraformHelm

Browse more by category

Show 171 moreSite Reliability EngineerShow 657 moreRecruiterShow 34 moreELKShow 2,794 moreAWSShow 832 moreTerraformShow 182 moreDatadogShow 244 moreGrafanaShow 1,535 moreKubernetesShow 1,687 moreGoShow 147 moreAnsible
Privacy·Terms··Contact·FAQ·Wagey on X