Hands-on experience using Slurm in production from a user’s perspective, including submitting and debugging workloads with sbatch, srun, squeue, and sinfo
Strong proficiency in Go, with experience building production-grade Kubernetes operators, controllers, CRDs, and reconciliation loops
Experience preserving traditional Slurm cluster behavior while running the underlying infrastructure on Kubernetes
Experience diagnosing performance and reliability issues across GPUs, schedulers, hardware, high-performance networks, and distributed storage systems
A product mindset and strong customer empathy, treating Slurm as a customer-facing platform rather than simply another system daemon
Excellent communication skills and the ability to take end-to-end ownership of complex distributed-system challenges
Experience operating large-scale HPC or GPU clusters for external customers
Experience with PyTorch distributed training and other large-scale AI/ML frameworks
Experience with InfiniBand, RoCE, RDMA, GPUDirect, Lustre, WEKA, Ceph, or similar high-performance infrastructure
Experience building unified job-submission workflows across Kubernetes and Slurm
Experience in GPU-cloud or HPC product engineering environments
Contributions to Slurm, Kubernetes, Soperator, or other cloud-native and HPC open-source projects
Additional Information