Education & Experience: 12+ YOE demonstrated ability to competently and independently perform responsibilities plus Bachelor’s or Master’s degree in Computer Science, Electrical Engineering, or a related technical field
Experience building and deploying automated integration testing for an AI Cloud Environment, ranging from low-level Linux Systems up to Distributed Control Planes
Working knowledge of the modern infrastructure stack, including Kubernetes, Docker, Terraform, and Postgres
CI/CD & Gitlab: Intimate knowledge of CI/CD pipelines and Gitlab Tooling to enable stable infrastructure releases across multiple datacenters
Configuration Management: Previous experience with at least 1-2 configuration management systems, including Ansible, Puppet, Chef, or SaltStack
Automation & Scripting: Advanced proficiency in Python and/or Bash for automating complex cluster-wide test scenarios
System Internals: Knowledge of Linux kernel internals, specifically PCIe topology, VFIO, and memory management (HugePages, IOMMU)
Distributed GPU Ecosystems: Familiarity with NVIDIA (CUDA/NCCL) and/or AMD (ROCm/RCCL) stacks in a multi-node context
Networking Knowledge: Strong understanding of RDMA, RoCE, and InfiniBand protocols and their implementation in virtualized systems
Experience with MNNVL (Multi-Node NVLink) or specialized AI fabric architectures
Familiarity with hardware-level debugging tools and performance profilers (e.g., NVIDIA Nsight, AMD Omniperf)
Knowledge of containerized orchestration for GPUs (e.g., Kubernetes with specialized device plugins)