runpod - Datacenter Infrastructure Specialist
Requirements
• Professional Background: 3–5 years of experience in infrastructure operations, systems reliability, or datacenter engineering. • Datacenter Networking: Strong proficiency in standard datacenter networking and performance troubleshooting. Exposure to RDMA, InfiniBand, or RoCE is highly preferred. • GPU & AI Stack: Hands-on experience with the NVIDIA Software Stack (driver installation, performance utilities) and an understanding of multi-node performance tuning. • Systems & Diagnostics: Solid Linux system administration skills and experience with containerization (Docker). You are comfortable performing system-level troubleshooting and performance tuning at the kernel and hardware interface layers. • Effective Communication: Clear written and verbal communication skills. You can explain hardware or networking issues to both technical partners and internal leadership. • Operational Flexibility: As our global fleet scales, this role may require participating in an on-call rotation in the future. • Strategic Problem-Solver: You are detail-oriented and proactive when it comes to identifying potential failures before they impact customers. • Startup Experience: Experience working in a fast-paced environment where you have contributed to building operational workflows. • HPC Exposure: Experience managing or optimizing bare-metal High-Performance Computing environments at massive scale. • Observability Tools: Experience with Grafana, Prometheus, or Datadog to monitor system health. • Automation: Proficiency in Python, Go (Golang), or Bash to automate repetitive infrastructure tasks and interface with internal APIs. • What You’ll Receive: • The competitive base pay for this position ranges from . This salary range may be inclusive of several career levels at Runpod and will be narrowed during the interview process based on a number of factors, including the candidate’s experience, qualifications, and location. • Meaningful equity in a fast-growing AI infra company — everyone on the team receives stock options — your impact drives our growth, and you share in the upside. • Generous medical, dental & vision plans — we cover 100% for all employees and partial for dependents. • Flexible PTO — take the time you need to recharge. • Most roles are remote work first with inclusive, collaborative teams utilizing Slack as the main form of internal communication. • Join a passionate team on the cutting edge of AI infrastructure — where culture, learning, and ownership are at the heart of how we scale. • $1,200 Home Office & Equipment Stipend — We set you up for success from day one with gear and support to create your ideal workspace.
Responsibilities
• Hardware Validation & Benchmarking: Assist in validating new hardware, ensuring partner deployments meet Runpod’s specifications for distributed AI/ML workloads. • Uptime & SLA Enforcement: Monitor fleet health to identify performance degradation. You will help audit downtime and provide the technical data needed to protect customer SLAs. • AI-Driven Operations: We operate with an AI-first mindset, powering our operations with the technology we host. You will work with LLMs and AI agents to help automate network triage and generate dynamic runbooks for our fleet. • Incident Support: Coordinate technical incident communications with clear updates, acting as a steady hand that translates outages into actionable resolutions. • Partner Technical Support: Support the growth of our infrastructure partners
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT