5-6+ years of experience in hardware systems engineering, platform engineering, performance engineering, ML systems engineering, infrastructure engineering, or related areas
Hands-on experience with large-scale GPU or accelerated computing infrastructure for AI/ML or HPC workloads
Hands-on experience with distributed training and/or inference workloads at scale, including parallelism strategies and performance tuning across the hardware/software stack
Experience with workload benchmarking, performance profiling, and system performance optimization across hardware and software layers
Strong understanding of modern server and accelerator architectures, including CPU, GPU, memory, storage, networking, and high-speed interconnects such as PCIe, InfiniBand, or NVLink
Hands-on experience with system bring-up, validation, performance characterization, and root-cause analysis of complex hardware/software issues
Experience developing automation, testing, diagnostics, or data-analysis frameworks using Python, Shell, or similar languages
Ability to analyze system behavior using telemetry, benchmarks, profiling tools, and other quantitative data
Experience working across multiple engineering disciplines, including hardware, firmware, software, networking, and infrastructure teams
Strong analytical and problem-solving skills with the ability to operate effectively in ambiguous and rapidly evolving environments
Excellent technical communication skills and experience collaborating with internal engineering teams, customers, and external technology partners
Bachelor's or Master's degree in Electrical Engineering, Computer Engineering, Computer Science, or equivalent experience
Experience influencing hardware or system configuration decisions based on workload performance data (e.g: HW/SW co-design, platform tuning studies)
Deep experience with RDMA, RoCE, CXL, NVLink or fabric-level performance analysis
Experience with inference serving frameworks, training frameworks, or ML compiler/runtime stacks
Familiarity with both x86 and ARM-based server platforms
Experience building observability, diagnostics, or fleet-level performance and reliability systems
Experience introducing new compute technologies into production cloud or large-scale datacenter environments
Understanding of infrastructure efficiency, power, cooling, performance-per-dollar, or total cost of ownership considerations
Background in sustainable or energy-efficient hardware design practices
Advanced certifications or coursework in AI/HPC hardware systems