Own day to day operations of electrical and mechanical systems including UPS, generators, switchgear, PDUs, chillers, CRAH/CRAC units, and cooling infrastructure
Ensure high availability and uptime of all critical systems supporting data center operations
Lead incident response for facility related events, including root cause analysis and corrective actions
Monitor system performance via BMS/DCIM and drive improvements in reliability and efficiency
Oversee preventive and corrective maintenance programs for all facility systems
Manage and hold accountable third party vendors and service providers
Ensure all maintenance activities follow SOPs, EOPs, and MOPs
Drive standardization and continuous improvement of operational procedures
Partner with engineering and operations teams on capacity planning and infrastructure scaling
Monitor and improve PUE and overall energy efficiency
Identify and implement reliability and sustainability improvements across facilities
Support high density environments, including air and liquid cooling strategies aligned with modern AI workloads
Ensure compliance with local regulations, safety standards, and industry best practices
Maintain strong adherence to HSE (Health, Safety, Environmental) standards
Lead audits, inspections, and documentation for operational readiness
Act as the primary point of contact for facility related regulatory interactions
Support commissioning, testing, and handover of new or expanded infrastructure
Participate in FAT/SAT and system validation for critical equipment
Ensure smooth transition from construction to steady state operations
Validate that systems are operationally ready with proper documentation and procedures
Partner closely with
Data Center Operations teams
Infrastructure & deployment engineers Network and hardware teams
Act as the bridge between facilities and IT infrastructure, ensuring both layers operate seamlessly (a key distinction in Nebius environments)
Support broader operational goals around scalability, reliability, and standardization