• 10+ years of engineering leadership in large-scale distributed systems, infrastructure, or technical supply chain, with a track record leading compute platform strategy at a frontier AI lab, hyperscaler, or major autonomy program.
• Deep technical and commercial fluency in cluster topology, high-speed interconnects (InfiniBand/RoCE), large-scale data systems, and the economics of distributed training.
• Direct operational oversight of 10k+ accelerator environments in production.
• Experience orchestrating capital or infrastructure for training runs at the 100B-parameter or 100k-GPU-day scale.
• Familiarity with the capacity and latency demands of edge-to-cloud inference and real-time autonomous systems.