Lead planning and execution of major AI infrastructure initiatives spanning training pipelines, data systems, model evaluation, and inference/serving
Build structures that keep teams aligned: scopes, goals, requirements, timelines, risks, and success metrics
Partner with engineering, research, and product to translate model and product needs into infrastructure roadmaps and priorities
Drive cross-functional accountability and communication across teams working on tightly coupled systems
Track key infrastructure metrics (e.g., reliability, latency, throughput, cost efficiency) and define reporting that surfaces progress and risk
Identify bottlenecks in infrastructure workflows and lead efforts to improve tooling, automation, and developer velocity
Support capacity planning and resource allocation to ensure infrastructure scales with model and product growth
Develop repeatable frameworks and operational patterns that improve execution quality across infrastructure programs
Serve as a strategic partner to leaders on prioritization, sequencing, and tradeoffs across infrastructure investments
Work closely with AI cloud providers and serve as the primary bridge between internal engineering teams and external partnership efforts
Technical understanding of GPU clusters, AI accelerator chips, and cloud providers to effectively evaluate requirements and guide infrastructure decisions