Experience with virtualization and container (Docker, Kubernetes) technologies
Experience with neoclouds/GPU cloud providers
Flexible availability for potential shifts outside of normal working hours/weekends
Experience with high performance storage systems
Familiarity with infrastructure-as-code tools (Terraform, Ansible, etc.)
Experience with Nvidia GPUs and Infiniband
Salary Range Information
3+ years of hands-on HPC experience in an administration, support, or engineering role
Very strong understanding and experience supporting Linux in a system administration role
Proven experience in HPC environments, showcasing your expertise in Linux cluster administration, with strong preference for Kubernetes and/or Slurm for cluster orchestration
Strong coding ability and CI/CD experience, with a track record of using AI-assisted tools to move fast
Proficiency with monitoring/logging tools (Prometheus, Grafana, Datadog)
Strong skills in log analysis, debugging kernel-level issues, and performance profiling
Experience with CUDA, NCCL, NVLink, GPUDirect RDMA
Experience with high throughput networking technologies(IB/RoCE)
Knowledge of distributed AI/ML or HPC workloads
Knowledge of TCP/IP, VPN, and firewalls in cloud environments
Ability to work independently and mentor junior support engineers