TensorWave is seeking an experienced Technical Program Manager with a strong data center operations background to lead and scale the operational programs that keep our next-generation AI infrastructure running at peak performance
In this role, you’ll own the full lifecycle of data center operations programs: from hardware deployment and capacity management through incident response, change management, and continuous reliability improvement. You’ll be the operational spine connecting facilities engineers, network, hardware, DevOps, and SRE teams, and executive leadership to ensure our AMD-powered AI clusters deliver on customer commitments at scale, with zero tolerance for preventable downtime
This is a high-visibility, high-impact role for someone who brings operational discipline to complex, fast-moving environments and maintains clear communication and structured execution under pressure
Own end-to-end program management for data center operations across multiple sites, covering hardware lifecycle, capacity planning, change management, incident response, and operational readiness
Serve as the primary coordination point across facilities, networking, hardware, and software teams: driving accountability to operational SLAs, program schedules, and customer commitments
Define and track program milestones, critical path dependencies, and resource requirements across concurrent multi-site operational programs
Translate operational status and risk into clear reporting for engineering, product, and executive leadership audiences
Identify, escalate, and mitigate risks to site reliability, capacity availability, and customer-facing uptime before they become incidents
Coordinate hardware deployment, node lifecycle, and maintenance sequencing across sites in alignment with capacity and customer commitments
Partner with network, power, facilities, and infrastructure engineering teams to drive operational readiness for high-density GPU compute clusters
Own post-incident retrospectives and corrective action tracking, driving lessons learned into durable process improvements
Maintain program documentation including operational runbooks, risk registers, vendor trackers, capacity plans, and change management records