Minimum 5 years of hands-on systems engineering experience across Linux, server hardware, firmware, PCIe, and GPU or high-performance computing platforms
Extensive Linux experience, particularly using Linux as an environment for hardware debugging, system-level troubleshooting, and platform investigation
Knowledge of the Linux kernel and experience with kernel-level debugging or troubleshooting
Strong knowledge of modern server architecture, particularly in high-performance, GPU-based environments
Strong knowledge of NVIDIA GPU platforms and diagnostic tooling, including nvidia-smi, XID and SXID analysis, GSP and driver behaviour, NVLink, NVSwitch, Fabric Manager, DCGM diagnostics, and NCCL testing
Deep knowledge of PCIe, including topology, enumeration, root complexes, endpoints, switches, bridges, retimers, link speed and width, AER and DPC errors, completion timeouts, and link failures. Candidates should understand how protocol-level errors can relate to physical-layer or signal-integrity problems
Experience benchmarking systems for performance, stability, power efficiency, thermal behaviour, and workload characteristics
In depth understanding of firmware interactions across BIOS/UEFI, BMC, CPLD/FPGA, GPU firmware, NIC firmware, and other platform components
Experience with low-level hardware communication and debugging using I²C, SMBus, PMBus, register maps, bit masks, byte- and word-level data, and device datasheets
Demonstrated ability to troubleshoot complex hardware, software, and networking issues
Experience with deep problem investigation, root cause analysis, and performance optimization in cloud or high-performance computing environments
Strong analytical and problem-solving skills with a performance-first mindset
Scripting and automation using various languages (Python\Go)