Experience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)
Experience building and/or operating HPC resources
Depth in the NVIDIA hardware and firmware ecosystem
Experience with data center power and thermal design
Background in chaos engineering or similar reliability testing methodologies
Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)
Salary Range Information
7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization
Strong understanding of Linux-based systems in a distributed environment
Are experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environments
Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling
Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)
Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)
Have excellent problem-solving and troubleshooting skills and an innate attention to detail
Passion for continuous improvement and innovation