wagey.ggwagey.gg
30,547  jobs30,547  jobs
Browse Tech JobsCompaniesFeaturesPricingFAQs
Log InGet Started Free
Jobs(30,547)/Director of Engineering Role(176)/runpod (10) - Director of Infrastructure Engineering
runpod

runpod - Director of Infrastructure Engineering

Remote - USA$225k - $325k+ Equity1w ago
RemoteDirectorNAMaterialsNonprofitDirector of EngineeringInfrastructure EngineerStakeholder ManagementDocumentationKubernetesTerraformAnsibleSlackProgram Management

Requirements

• Engineering Leadership Experience: 7+ years leading software, infrastructure, SRE, or networking teams, including managing managers and multiple squads, with a proven record of scaling high-availability cloud environments. • Deep Infrastructure Expertise: 8+ years building and operating large-scale distributed systems, bare-metal infrastructure, or public/private cloud platforms. • HPC & Advanced Networking: Proven hands-on background or strong architectural understanding of ultra-low latency networking. Deep familiarity with InfiniBand and/or RoCE, spine-leaf architectures, and global WAN routing protocols (BGP). • Storage Systems Knowledge: Experience building, operating, or tuning high-performance distributed storage systems and parallel file systems (e.g., Ceph, Lustre, Weka, NVMe-oF) capable of handling heavy AI/ML I/O loads. • SRE / DevOps Culture: Strong foundation in reliability engineering, infrastructure-as-code (Terraform, Ansible), container orchestration (Kubernetes), and modern observability stacks. • Remote-First Operating Excellence: Experience building culture, accountability, and momentum across distributed technical teams. • Communication & Collaboration: Clear written and verbal communication, strong stakeholder management, and calm, decisive leadership during high-stakes operational incidents. • Background Check: Successful completion of a background check. • Direct experience architecting and operating infrastructure specifically optimized for massive GPU clusters and AI/ML workloads. • Deep understanding of hardware architectures, GPU interconnects (NVLink), and datacenter topology. • Track record of scaling infrastructure teams in hyper-growth startup environments. • Open-source contributions or active recognition within the infrastructure, networking, or Kubernetes communities. • What You’ll Receive: • The competitive base pay for this position ranges from ($225,000 - $325,000). This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location. • Meaningful equity in a fast-growing company- everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans. • Flexible PTO- take the time you need to recharge. • Most roles are remote work first with an inclusive, collaborative teams utilizing slack as the main form of internal communication. • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. • $1,200 Home Office & Equipment Stipend- We set you up for success from day one with gear and support to create your ideal workspace.

Responsibilities

• Own Core Infrastructure & SRE: Lead multiple engineering teams responsible for Site Reliability Engineering, networking, and storage. Establish rigorous SRE practices, driving SLA/SLO definitions, incident response, observability, and automated remediation. • Architect HPC & Global Networking: Oversee the design, scaling, and operation of Runpod’s global network backbone, as well as ultra-low-latency HPC cluster networks. Drive the implementation and optimization of InfiniBand and RDMA over Converged Ethernet (RoCE) to support massive, multi-node GPU training workloads. • Drive Storage Engine Innovation: Direct the architecture and performance tuning of highly scalable, distributed storage systems. Ensure our storage engines can deliver the massive IOPS and throughput required to keep high-end GPUs fed with data during deep learning tasks. • Build a High-Output Org: Hire, mentor, and grow highly technical engineering managers and senior ICs (network architects, systems engineers, SREs). Create a culture of ownership, operational excellence, and craft in a remote-first environment. • Translate Scale into Strategy: Partner with Program Management and Product to forecast capacity requirements, shape technical roadmaps, and convert massive scale challenges into clear technical scopes, milestones, and measurable outcomes. • Continuously Improve Systems & Flow: Drive measurable improvements in infrastructure reliability and delivery metrics, such as deployment frequency, MTTR (Mean Time To Recovery), infrastructure as code (IaC) coverage, and system uptime. • Architectural Stewardship: Provide architectural oversight for bare-metal provisioning, virtualization layers, network fabrics, and storage clusters, ensuring seamless scalability without becoming a bottleneck for your teams. • Cross-Functional Partnership: Coordinate cleanly with product delivery and platform teams to ensure the infrastructure primitives they rely on are robust, well-documented, and highly available.

Apply in one click

Upload My Resume

Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT

Apply in One Click
Apply in One Click

Similar roles

EthycaEthyca - Director, Mission Engineering Group3mo ago
·Remote - New York, United States·$160k - $240k/year + Equity
RemoteNADirectorPaymentsDirector of EngineeringDocumentationTeam ManagementProgram ManagementStripeSnowflake
runpodrunpod - Director, Forward Deployed Engineering & Technical Support1w ago
·Remote - USA - Hybrid·$300k - $300k/year + Equity
In OfficeNADirectorCloud ComputingArtificial IntelligenceDirector of EngineeringPythonDockerKubernetesZendeskNotionClaudeGeminiLinearDocumentationSlackCSATAccount ManagementSales Strategy
emergenceemergence - Director of IT1w ago
·Remote - USA·$140k - $175k/year
RemoteNADirectorData AnalyticsDirector of EngineeringDocumentationCloseDue DiligenceGoogle WorkspaceMicrosoft 365Stakeholder ManagementBack-end
BugcrowdBugcrowd - Director, Technical Solutions1w ago
·Remote - USA·$147k - $184k/year
RemoteNADirectorCybersecurityArchitectureDirector of EngineeringTeam ManagementKPI TrackingProcess OptimizationProgram ManagementCRM ManagementDocumentationGovernance
RailwayRailway - Senior Infra Engineer: Baremetal Orchestration2mo ago
·San Francisco, California , USA·Equity
RemoteNASeniorMaterialsJunior EngineerInfrastructure EngineerAnsibleDocumentationRailwayTerraformRust
Orcrist TechnologiesOrcrist Technologies - (GPU) Infrastructure Engineer1mo ago
·Remote - Germany
RemoteEMEASeniorArtificial IntelligenceMaterialsInfrastructure EngineerTerraformAnsibleCUDAKubernetesPulumi
Clover HealthClover Health - Director, Engineering (Interventions)1mo ago
·Remote - USA·$223k - $223k/year + Equity
RemoteNADirectorDiagnosticsArtificial IntelligenceDirector of EngineeringMentoringResource AllocationStakeholder Management
Blackpoint CyberBlackpoint Cyber - Director of SRE2d ago
·Canada·$118k - $151k/year
In OfficeNADirectorLife InsuranceInsuranceDirector of EngineeringAWSAzureGCPTeam ManagementCoachingTerraformKubernetesPrometheusGrafanaDatadogSplunkPulumiLokiProcurement
angiangi - Senior Infrastructure Engineer4w ago
·Remote - Denver, Colorado, United States·$165k - $216k/year + Equity
RemoteNASeniorCloud ComputingInfrastructure EngineerAWSKubernetesTerraformPythonDocker

Browse more by category

Show 176 moreDirector of EngineeringShow 191 moreInfrastructure EngineerShow 861 moreStakeholder ManagementShow 4,823 moreDocumentationShow 1,520 moreKubernetesShow 829 moreTerraformShow 146 moreAnsibleShow 431 moreSlackShow 803 moreProgram Management
Privacy·Terms··Contact·FAQ·Wagey on X