cerebras - Sr./Staff TPM - Inference Capacity
Requirements
• 5+ years of TPM, technical program management, or product operations experience in cloud infrastructure, large-scale ML serving, or hyperscaler capacity planning • Experience leading large cross-functional programs involving Engineering, Product, and Operations • Comfort with the inference serving stack: model replicas, batching, prefill/decode, KV cache, accelerator scheduling • Strong data fluency: SQL, Grafana, basic Python or Flux to pull your own numbers without waiting for an analyst • Track record of running a recurring cross-functional ritual involving senior engineers and LT • Direct experience with AI accelerator fleet operations such as Habana, TPU pods, Inferentia, Trainium
Responsibilities
• Run weekly capacity planning and daily capacity and deployment tracking with Engineering, product and operations team. Own fleet utilization reporting and forecasting • Drive capacity planning for new customer deployments and major model launches • Drive continuous improvement and stakeholder adoption of new capacity management platform • Drive org level strategic initiatives related to capacity expansion, improving fleet efficiency and maximizing effective utilization of available systems • Lead planning around major infrastructure events including but not limited to new customer commits, new model releases, change to DC/cluster architecture, etc. that impacts capacity and fleet utilization. Update capacity plans and forecasts accordingly. • Maintain Jira EPICs and Confluence pages related to capacity planning, reporting and change management to ensure execution transparency across teams
Benefits
• People who are serious about software make their own hardware. At Cerebras, we have built a breakthrough architecture that is unlocking new opportunities for the AI industry. With dozens of model releases and rapid growth, we’ve reached an inflection point in our business. Members of our team tell us there are five main reasons they joined Cerebras: • 1. Build a breakthrough AI platform beyond the constraints of the GPU. • 2. Publish and open source their cutting-edge AI research. • 3. Work on one of the fastest AI supercomputers in the world. • 4. Enjoy job stability with startup vitality. • 5. Our simple, non-corporate work culture that respects individual beliefs. • Find out more about what it's like to work at Cerebras here https://www.cerebras.ai/join-us!
Apply in one click
Upload My Resume
Drop here or click to browse · Tap to choose · PDF, DOCX, DOC, RTF, TXT