lambda - HPC Support Engineer
Requirements
• Experience with virtualization and container (Docker, Kubernetes) technologies. • Experience with neoclouds/GPU cloud providers. • Flexible availability for potential shifts outside of normal working hours/weekends. • Experience with high performance storage systems. • Familiarity with infrastructure-as-code tools (Terraform, Ansible, etc.) • Experience with Nvidia GPUs and Infiniband.
Responsibilities
• Serve as a senior technical escalation point, troubleshooting the hardest infrastructure and platform issues down to the hardware, driver, or kernel level when needed • Quickly and accurately distinguish between hardware failures, driver issues, kernel-level problems, and customer workload misconfiguration, so issues get resolved correctly the first time • Proactively identify process, tooling, and documentation gaps, and go fix them, not just wait for them to be assigned • Use AI tools effectively to build scripts, automations, or small internal tools that close real operational gaps (no professional development background required) • Perform root-cause analysis across distributed systems, clusters, and GPU infrastructure • Craft clear documentation of solutions and contribute to evolving support procedures • Collaborate closely with engineering teams to turn recurring customer pain points into permanent fixes • Take escalations from peers while training and mentoring them in the process • Participate in a rotating on-call schedule, owning major incidents and major customer issues • Be ready to roll up your sleeves and pitch in wherever needed, especially during fast, high-volume deployments • 3+ years of hands-on HPC experience in an administration, support, or engineering role. • Very strong understanding and experience supporting Linux in a system administration role. • Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for Kubernetes and/or Slurm for cluster orchestration. • Strong coding ability and CI/CD experience, with a track record of using AI-assisted tools to move fast. • Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog). • Strong skills in log analysis, debugging kernel-level issues, and performance profiling. • Experience with CUDA, NCCL, NVLink, GPUDirect RDMA. • Experience with high throughput networking technologies(IB/RoCE). • Knowledge of distributed AI/ML or HPC workloads. • Knowledge of TCP/IP, VPN, and firewalls in cloud environments. • Ability to work independently and mentor junior support engineers.
Benefits
• This is a salaried exempt role. The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description.
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT