wagey.ggwagey.gg
31,365  jobs31,365  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(31,365)/Site Reliability Engineer Role(176)/Andromeda Cluster (8) - Senior Site Reliability Engineer - AI Infrastructure
Andromeda Cluster

Andromeda Cluster - Senior Site Reliability Engineer - AI Infrastructure

Remote - San Francisco, California , United States4mo ago
RemoteSeniorNASite Reliability EngineerBashGoHelmPythonTerraform

Requirements

• GPU Systems Expertise: Deep, hands-on experience operating large-scale GPU clusters (NVIDIA A100/H100/B200 or equivalent). You understand GPU memory hierarchies, ECC behavior, thermal throttling, and hardware failure modes from direct experience not documentation. • High-Performance Networking: Production experience with InfiniBand, RoCE, or NVLink fabrics in the context of distributed training. You can diagnose why an all-reduce is slow, identify a degraded link in a fat-tree topology, and reason about congestion control at scale. • Distributed Training & ML Frameworks: Working knowledge of how large training jobs actually run — NCCL, CUDA, PyTorch distributed, DeepSpeed, Megatron, FSDP, or similar. You don't need to write the models, but you need to understand what's happening at the systems level when a 1,000-GPU training run stalls. • Linux & Systems Internals: Expert-level Linux knowledge: kernel tuning, driver management (NVIDIA drivers, CUDA toolkit), cgroup/namespace internals, performance profiling at the syscall and hardware level. • Kubernetes & Orchestration: Strong experience running Kubernetes in production with GPU workloads, including device plugins, topology-aware scheduling, multi-cluster federation, and custom operators. Experience with Slurm or other HPC schedulers is equally valued. • Automation & Software Engineering: Strong engineering skills in Python, Go, or Bash. You build production-grade tools and services, not just scripts. Infrastructure-as-Code proficiency (Terraform, Helm, Ansible, or equivalent). • Observability & Monitoring: Hands-on experience building monitoring and alerting for GPU infrastructure, not just Prometheus/Grafana basics, but GPU-specific telemetry (DCGM, nvidia-smi, fabric manager metrics) integrated into actionable dashboards. • Incident Management: Proven track record leading incident response for complex distributed systems where the failure could be in hardware, firmware, networking, drivers, orchestration, or application code and you need to narrow it down fast. • Strong Candidates May Have • Distributed Storage: Experience with high-performance parallel file systems (VAST, Weka, Lustre, GPFS) and the checkpoint I/O and data-loading bottlenecks that come with large training runs. • Training Optimization: Experience profiling and optimizing distributed training performance: identifying stragglers, tuning collective communication strategies, improving MFU (Model FLOPs Utilization), and reducing idle GPU time across large runs. • Cluster Buildout & Hardware: Experience involved in physical cluster design - rack layout, power/cooling constraints, network topology design, and hardware validation/burn-in at scale. • Team Leadership: Experience leading or mentoring a team of infrastructure engineers. We're growing and need people who raise the bar for everyone around them.

Benefits

• This is a high-impact, senior builder’s role. You’ll have significant ownership and autonomy to shape how our systems run at a foundational level, working directly with customers and providers while architecting the infrastructure backbone for reliable, scalable AI compute. You’ll influence technical direction and help define what world-class AI infrastructure operations look like.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

Bedrock Ocean ExplorationBedrock Ocean Exploration - Senior Site Reliability Engineer, Robotics & Cloud Infrastructure1mo ago
·Brooklyn, New York, USA·$164k - $220k/year + Equity
In OfficeNASeniorCloud ComputingRoboticsSite Reliability EngineerTeam ManagementCustomer OnboardingBashPythonGo
AmwellAmwell - Senior Site Reliability Engineer1mo ago
·US - Remote - Hybrid·$129k - $140k/year + Equity
In OfficeNASeniorMental HealthCloud ComputingSite Reliability EngineerBashPythonPuppetTerraformAnsible
SecurityScorecardSecurityScorecard - Senior Site Reliability Engineer3w ago
·Remote - USA·$152k - $195k/year + Equity
RemoteNASeniorArtificial IntelligenceSite Reliability EngineerBashGoPythonKubernetesJenkinsTerraformHelmPrometheusGoogle GKEGovernancePulumiGrafanaDatadogKafkaMLOps
BranchBranch - Senior Site Reliability Engineer (SRE)1w ago
·Remote - USA·$364k+/year
RemoteNASeniorCloud ComputingSite Reliability EngineerJavaSpring BootTerraformDockerGoKubernetesBashPythonGoogle Pub/SubRedisPrometheusMySQLGrafanaGCP
Epic Kids Inc.Epic Kids Inc. - Senior Site Reliability Engineer3w ago
·Remote - USA
RemoteNASeniorCloud ComputingSite Reliability EngineerBashPythonMandarinGCPDockerKubernetesHelmJenkinsTerraformNew RelicGoogle GKEAirflowDagsterCFPSAFe
PinterestPinterest - Site Reliability Engineer II, tvScientific1mo ago
·San Francisco, California, United States·$114k - $114k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerBashPythonAWSKubernetesTerraform
Chainlink LabsChainlink Labs - Site Reliability Engineer II5mo ago
·Remote - Canada, United States, Brazil...
RemoteNAMidCryptocurrencyCloud ComputingSite Reliability EngineerGoShellPythonKubernetesTerraform
runpodrunpod - Site Reliability Engineer2w ago
·Remote - USA·$150k - $200k/year + Equity
RemoteNASeniorSite Reliability EngineerLinuxPerformance ReviewsReportingSlackPrometheusGrafanaPythonBashGo
FlockFlock - Site Reliability Engineer III, DevEx1w ago
·Remote - United States·$140k - $165k/year + Equity
RemoteNAMidCloud ComputingSite Reliability EngineerGoTypeScriptKubernetesTerraformHelmAWS

Browse more by category

Show 176 moreSite Reliability EngineerShow 358 moreBashShow 1,805 moreGoShow 135 moreHelmShow 5,064 morePythonShow 921 moreTerraform
Privacy·Terms··Contact·FAQ·Wagey on X