tensorwave - Senior Staff Infrastructure Engineer – Kubernetes Platform
Requirements
• 7+ years of experience in infrastructure, platform engineering, or distributed systems • Deep experience operating Kubernetes at scale in production environments • Experience in CSP, hyperscale, or equivalent large-scale environments strongly preferred • Proven experience scaling Kubernetes across: • Multiple clusters • Multiple regions or data centers • Strong understanding of Kubernetes internals: • Controller manager • Control plane architectures • Multi-tenant cluster models • Technical Depth • Strong Linux systems expertise • Deep troubleshooting ability across: • Container runtime • Networking stack • Experience with CNI plugins (Cilium preferred) • Strong understanding of: • Networking and traffic patterns • Resource isolation and scheduling • Experience with virtual cluster technologies (vcluster, Kamaji, or similar) • Experience supporting GPU workloads in Kubernetes • Familiarity with: • NUMA-aware scheduling • Topology-aware workloads • Awareness of RDMA and high-throughput networking environments • Experience with observability platforms (Prometheus, Grafana, etc.)
Responsibilities
• Platform Architecture & Strategy • Design and evolve Kubernetes control plane architecture across regions • Define and implement multi-tenant cluster models, including shared control planes, virtual cluster approaches (e.g., vcluster, Kamaji) • Drive transition from standalone clusters to regionally managed platform models • Define standards for isolation boundaries, resource segmentation, policy enforcement • Platform Ownership & Operations • Own the reliability and behavior of Kubernetes platforms in production • Participate in on-call rotation and lead incident response • Diagnose and resolve control plane instability, API server saturation, scheduling and resource contention issues • Ensure consistent lifecycle management across clusters - provisioning, upgrades, scaling • Multi-Region Scaling • Design and implement strategies for regional scaling, multi-data center cluster deployments • Ensure consistent behavior and reliability across environments • Define cluster topology and failure domain strategies • Networking & Data Plane Integration • Design ingress and egress architectures at cluster level and regional level • Troubleshoot and optimize pod-to-pod networking, north-south traffic flows, CNI behavior (Cilium preferred) • Collaborate with network engineering on high-performance networking integration • Observability & Reliability • Improve observability across control plane components, cluster health and performance • Define and implement resilience strategies aligned with platform goals • Lead root cause analysis for production incidents • Cross-Team Collaboration • Work closely with DevOps engineers (automation and CI/CD) and Infrastructure teams (compute, storage, networking) • Align Kubernetes platform design with underlying infrastructure capabilities
Benefits
• Stock Options • 100% paid Medical, Dental, and Vision insurance for Employees • Company Health Savings Account Contributions • 100% paid Short Term and Long Term Disability Insurance for Employees • Life and Voluntary Supplemental Insurance Options • Other Insurance Options, such as Pet & Legal Insurance • Various Supplementary Health Benefits, such as discounted Virtual Healthcare Appointments and Serious Illness Support • Flexible Spending Account • Employee Assistance Program • Paid Holidays • Parental Leave • Equal Employment Opportunity
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT