Fleet Engineering spans several teams. Depending on the team you join, your day-to-day will involve some combination of the following
Develop and Maintain Production Systems: Design, implement, and improve the software that powers fleet lifecycle management, machine configuration, and cluster state at scale
Automate Provisioning and Deployment: Build and enhance automation that takes clusters from logical design and racking through OS provisioning, configuration, validation, and customer hand-off
Support New Hardware and Site Bring-Up: Enable bring-up, validation, and production readiness for new server, accelerator, and network platforms, as well as new datacenter sites
Improve Machine Lifecycle Workflows: Refine bare metal provisioning, firmware and DPU updates, imaging, and system health monitoring across the fleet
Keep Fleet State Consistent and Healthy: Build systems that reconcile intended against actual configuration, catch drift before it causes deployment failures, and maintain production SLAs
Debug Hardware and Firmware Issues: Investigate failures across BIOS, BMC, firmware, DPUs, networking, storage, and boot flows
Collaborate Across Teams: Work closely with datacenter and deployment operations, networking, architecture, security, and product engineering teams to build scalable, maintainable solutions