5–8 years of professional software engineering experience in distributed systems, cloud infrastructure, or platform engineering
Strong experience building production systems in Go (Python or C++ a plus)
Solid understanding of Kubernetes fundamentals, APIs, controllers, and operating services in production
Experience working with scheduling, resource management, or quota-based systems
Proven ability to improve system reliability and performance using data and operational metrics
Comfortable owning services in production and participating in on-call rotations
Bachelor’s Degree in Computer Science, Engineering or related technical field
Experience with Kubernetes-native orchestration frameworks such as Kueue, Volcano, Ray, Kubeflow, or Argo Workflows
Familiarity with GPU-based workloads, ML training, or inference pipelines
Knowledge of scheduling concepts such as quota enforcement, pre-emption, and backfilling
Experience with reliability practices including SLOs, alerting, and incident response
Exposure to AI infrastructure, HPC, or large-scale distributed compute environments