We are building large-scale, high-performance infrastructure to power next-generation AI workloads. Our platform operates across multiple data centers and supports GPU-intensive environments with demanding requirements around performance, isolation, and scalability
We are looking for a Staff Infrastructure Engineer to lead the design and evolution of our virtualization platform. This role will own how we build, scale, and operate hypervisor infrastructure as we transition from traditional virtualization platforms toward a more flexible, CSP-aligned architecture based on KVM/QEMU and modern Linux primitives
This is a highly technical, hands-on role focused on solving complex systems problems at scale
Design and implement a scalable virtualization platform capable of supporting high-density compute and GPU workloads
Lead the evolution from existing platforms (e.g., Proxmox) toward KVM/QEMU-based architectures
Define standards for VM lifecycle management (provisioning, scheduling, migration), performance isolation and resource allocation, failure domains and resilience strategies
Optimize virtualization for high-performance workloads, including NUMA alignment, CPU pinning and scheduling, PCIe topology awareness, GPU passthrough and device assignment
Partner closely with networking and storage teams to integrate high-throughput, networking (e.g., SR-IOV, RDMA), distributed and local storage systems
Build and improve automation for hypervisor deployment and configuration, image pipelines, cluster scaling and lifecycle management
Troubleshoot deep system-level performance issues across compute, memory, storage, and network layers
Contribute to long-term platform architecture and infrastructure strategy