Experience operating at scale in data center, cloud infrastructure, or hyperscale environments balancing reliability investments against capacity goals, with an understanding of how operational decisions affect sellable capacity
Familiarity with reliability frameworks such as SLIs/SLOs, error budgets, incident management practices, and root cause analysis
Understanding of hardware failure and processes, including failure diagnosis, vendor coordination, and replacement lifecycle management
Background in observability, monitoring, or telemetry systems (e.g., Prometheus, Grafana, OpenTelemetry)
Experience with fleet lifecycle management (provisioning, firmware/OS updates, decommissioning) at scale
Bachelor's degree in Computer Engineering, Computer Science, or a related technical field, or equivalent practical experience
7+ years of technical program management experience in large-scale compute infrastructure, cloud, or platform environments
Experience operating in a rapidly growing or fast-scaling infrastructure environment, with programs designed to hold up under 2x, 5x, or greater fleet growth
Strong technical aptitude across infrastructure domains (compute, storage, networking, hardware, or SRE)
Demonstrated ability to use data and metrics to drive prioritization, execution, and decision-making
Excellent communication and stakeholder management skills, including executive-level reporting