wagey.ggwagey.gg
31,341  jobs31,341  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,341)/Site Reliability Engineer Role(176)/Bedrock Ocean Exploration (1) - Senior Site Reliability Engineer, Robotics & Cloud Infrastructure
Bedrock Ocean Exploration

Bedrock Ocean Exploration - Senior Site Reliability Engineer, Robotics & Cloud Infrastructure

Brooklyn, New York, USA$164k - $220k+ Equity2mo ago
In OfficeSeniorNACloud ComputingRoboticsSite Reliability EngineerTeam ManagementCustomer OnboardingBashPythonGo

Requirements

• 5+ years in an SRE, DevOps, or infrastructure engineering role running production systems with real uptime and on-call responsibilities, including senior-level ownership of reliability outcomes. • Experience implementing a scalable incident management and operational excellence mechanism that treats operators as customers, building processes and tooling that serve the people running operations day to day, not just the engineering team. • Strong automation instincts: comfortable scripting and building tooling in Python and/or Go and Bash, and using infrastructure-as-code (Terraform or equivalent). • Hands-on AWS experience across compute, storage, networking, and IAM, plus containerization and orchestration (Docker, Kubernetes or similar). • Working knowledge of Linux internals, networking, and observability tooling (Prometheus/Grafana or equivalents). • Comfort operating across environments that aren’t just cloud: embedded or edge compute, intermittent connectivity, and physical systems that fail in messy ways. • A reliability mindset: you instrument before you guess, you automate the second time you do something manually, and you write things down so the next person or the system can handle it without you. • Strong ownership and communication in a small, fast-moving team. • Experience with robotics or embedded systems: ROS / ROS 2, Jetson or similar edge compute, sensor integration. • Background supporting field operations, autonomous systems, or hardware-in-the-loop environments. • Familiarity with data pipelines and geospatial or large-binary data formats. • Experience standing up on-call practices and incident response from an early stage. • Some connection to the ocean: professional, academic, or personal. You’re excited to be around people who dive, sail, build, and explore offshore. • Active U.S. Secret security clearance or above.

Responsibilities

• Own reliability across the full path from vehicle to customer: AUV onboard compute (Jetson-class modules, ROS 2), topside/operator systems, cloud data pipelines, and the platform that delivers data products. • Build and extend infrastructure automation- provisioning, configuration management, deployment, and self-recovery- so that routine field operations and pipeline runs require minimal manual intervention. • Design and improve observability: metrics, logging, tracing, and alerting that give both robotics and data teams early, actionable signal across vehicle fleets and cloud services. • Drive down on-call burden by identifying and eliminating single points of failure, writing runbooks, and automating the manual steps that currently require tribal knowledge. • Participate in a shared on-call rotation covering both robotics-side and cloud-side incidents in 12-hour shifts spanning European and East Coast business hours; lead and contribute to blameless post-incident reviews. • Define and track reliability targets, availability, data yield, recovery time, tied to continuous-operations goals, and partner with robotics and data teams to meet them. • Manage cloud infrastructure on AWS (compute, storage, networking, IaC, cost, and security posture) for data processing and platform workloads. • Improve fleet- and vehicle-level configuration management, deployment safety, and rollback so changes reach the field reliably and predictably. • Experience implementing a scalable incident management and operational excellence mechanism that treats operators as customers, building processes and tooling that serve the people running operations day to day, not just the engineering team. • Strong automation instincts: comfortable scripting and building tooling in Python and/or Go and Bash, and using infrastructure-as-code (Terraform or equivalent). • Hands-on AWS experience across compute, storage, networking, and IAM, plus containerization and orchestration (Docker, Kubernetes or similar). • Working knowledge of Linux internals, networking, and observability tooling (Prometheus/Grafana or equivalents). • Comfort operating across environments that aren’t just cloud: embedded or edge compute, intermittent connectivity, and physical systems that fail in messy ways. • A reliability mindset: you instrument before you guess, you automate the second time you do something manually, and you write things down so the next person or the system can handle it without you. • Strong ownership and communication in a small, fast-moving team.

Benefits

• Our biggest operational goal depends on systems that stay up and data that stays valid for long, continuous stretches with a small team and a limited rotation. The reliability and automation you build directly determines whether we can run continuous campaigns at scale. This is high-leverage infrastructure work with a clear, measurable mission. • Not a Fit If… • Not a Fit If… • You prefer environments where cloud and hardware never mix. • You’d rather build tickets than eliminate them. • You’re not comfortable with on-call ownership on a small team. • You want to optimize existing systems, not build the reliability practice alongside the product. • $164,000–$220,000 base salary annually, depending on location. The upper end of the range reflects compensation in the New York, NY metro. In addition, we offer comprehensive employee benefits and equity. • Work Authorization • Work Authorization • Candidates must have legal authorization to work in the United States without visa sponsorship. Bedrock does not sponsor employment visas. • Due to the nature of our government and defense work, candidates must be eligible to obtain a U.S. Secret security clearance if requested. An active Secret or higher clearance is not required to apply, but candidates who hold one are strongly preferred.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

Andromeda ClusterAndromeda Cluster - Senior Site Reliability Engineer - AI Infrastructure5mo ago
·Remote - San Francisco, California , United States
RemoteNASeniorSite Reliability EngineerBashGoHelmPythonTerraform
Stack AVStack AV - Senior Site Reliability Engineer2mo ago
·Remote - Pittsburgh, PA or Remote
RemoteNASeniorCloud ComputingGovernmentSite Reliability EngineerBashPythonLinuxGCPAWS
synthesiasynthesia - Senior Site Reliability Enigneer2mo ago
·Remote - USA
RemoteNASeniorCloud ComputingSite Reliability EngineerTeam ManagementAWSKubernetesMongoDBPython
AmwellAmwell - Senior Site Reliability Engineer2mo ago
·US - Remote - Hybrid·$129k - $140k/year + Equity
In OfficeNASeniorMental HealthCloud ComputingSite Reliability EngineerBashPythonPuppetTerraformAnsible
AuthZedAuthZed - Sr. Site Reliability Engineer3mo ago
·Remote - USA·$150k - $195k/year + Equity
RemoteNASeniorInsuranceCloud ComputingSite Reliability EngineerJavaRubyPythonGoDocker
playonsportsplayonsports - PlayOn - Senior Site Reliability Engineer4mo ago
·Remote - USA *·Equity
RemoteNASeniorCloud ComputingSoftwareSite Reliability EngineerJavaC++GoChange ManagementPython
DittoDitto - Senior Site Reliability Engineer5mo ago
·Remote - USA·$156k - $288k/year + Equity
RemoteNASeniorCloud ComputingSoftwareSite Reliability EngineerGoRustC++JavaPython
ClickHouseClickHouse - Senior Site Reliability Engineer- Remote6mo ago
·Remote - USA·$208k - $208k/year + Equity
RemoteNASeniorCloud ComputingSite Reliability EngineerPythonGoAWSAzureSQL
Wikimedia FoundationWikimedia Foundation - Senior Site Reliability Engineer, Infrastructure Foundations4mo ago
·Remote - Anywhere·$31k - $31k/year
RemoteWWSeniorLogisticsNonprofitSite Reliability EngineerTeam ManagementGoRubyBashPython

Browse more by category

Show 176 moreSite Reliability EngineerShow 2,893 moreTeam ManagementShow 433 moreCustomer OnboardingShow 358 moreBashShow 5,059 morePythonShow 1,803 moreGo
Privacy·Terms··Contact·FAQ·Wagey on X