Crusoe Cloud operates GPU infrastructure across six production sites globally, with a fleet that spans SuperMicro, HPE, and next-generation ODM platforms as we scale. We're looking for a Staff Data Center Operations Engineer to serve as the senior technical operations resource for the SiteOps org — based at our Denver headquarters, with cross-site scope and travel authority across our full portfolio
This role is the bridge between Crusoe's distributed site teams and our OEM and ODM hardware partners. You'll own platform-level escalations that exceed site-level capability, drive hardware decisions at the org level, and serve as SiteOps' technical presence at headquarters — visible to engineering, procurement, and leadership in a way that a field-based role cannot be
You'll be hands-on when the situation calls for it, traveling to sites for complex escalations, new platform bring-ups, and deployment support. But your primary leverage is organizational: building the technical standards, OEM relationships, and institutional knowledge that keeps Crusoe's GPU fleet reliable across every site
Develop and maintain deep technical relationships with Crusoe's primary hardware partners — currently SuperMicro and HPE, with upcoming ODM’s as growing platforms — at the engineering and field escalation level
Serve as Crusoe's technical voice in OEM/ODM partner conversations, surfacing field observations, influencing hardware roadmaps, and driving platform improvements that benefit the full fleet
Build familiarity with new ODM platform architecture, tooling, and escalation processes as Crusoe expands its ODM footprint
Support vendor evaluations and new platform qualifications in partnership with SiteOps and engineering leadership
Own OEM platform technical knowledge at the SiteOps org level — escalation playbooks, failure pattern analysis, and OEM relationship inputs across all sites
Own the development and maintenance of platform-specific SOPs, runbooks, and field troubleshooting procedures for the SiteOps org, ensuring site teams have current, actionable documentation across all active hardware platforms
Design and deliver technical training for SiteOps technicians covering hardware architecture, platform-specific troubleshooting, and field procedures — both for new hire onboarding and ongoing skill development as the fleet and team evolve
Contribute to the technician certification program and technical leveling standards across the org
Support new site bring-up efforts providing platform readiness and deployment execution expertise