Strong understanding of Kubernetes architecture, including CNI, CSI, operators, ingress/gateway, and control plane components
Hands-on experience operating and troubleshooting production Kubernetes clusters
Strong Linux and networking troubleshooting skills, including DNS, routing, firewalling, TLS, MTU, connectivity and performance issues
Ability to develop automation and operational tooling using Python, Go, or Bash
Experience with Terraform, Ansible, or similar IaC/configuration management tools
Experience with VictoriaMetrics/Grafana or similar monitoring, alerting, and troubleshooting tools
Strong experience with Git-based workflows and CI/CD pipelines
Familiarity with Cluster API or similar Kubernetes cluster lifecycle management technologies
Hands-on operation or administration of Slurm clusters
Knowledge of Argo CD, GitOps workflows, Helm, or Helmfile
Background working with managed platforms, PaaS, or cloud services
Exposure to bare metal, GPU, HPC, or other high-performance computing environments
Familiarity with the NVIDIA GPU stack, RDMA/InfiniBand, or high-performance networking
Knowledge of OpenStack or similar cloud infrastructure platforms
Hands-on experience developing Kubernetes operators or controllers
Additional Information