Minimum 1-2 years of software engineering experience, shipping and maintaining code in large production systems, ideally with breadth across the stack
Experience debugging complex production issues - working through logs, metrics, and traces to root-cause problems in unfamiliar systems
Confidence owning ambiguous technical problems. This includes triaging, making decisions under uncertainty, and knowing when to pull in other engineers who own the underlying systems
Motivation beyond pure engineering. An interest in wanting to work directly with customers, understand their problems firsthand, and influence the product
Clear communication on complex technical topics, whether you're talking to a customer's engineers or their leadership
Genuine curiosity about AI inference and training, and a drive to become an expert in the infrastructure powering it
Willingness to respond to customers outside regular working hours and participate in an on-call rotation
Excitement about solving problems for some of the largest and fastest-growing companies in the world running mission-critical AI workloads
We don't expect any one person to cover all of this - the strongest candidates may spike in one or two of the following
Depth in a core infrastructure domain such as storage systems or networking (anywhere from the cloud layer to cluster interconnects like InfiniBand and RoCE)
Experience operating distributed compute platforms like Kubernetes, Slurm, or Ray, especially for GPU workloads
A detailed understanding of LLM architectures and modern inference engines like vLLM, TensorRT-LLM, or SGLang
The ability to profile and optimize GPU workloads, in training or serving
Hands-on experience with post-training techniques like SFT and RL, or a broader deep learning background plus fluency in a tensor computation library like PyTorch or JAX
Operational depth - running on-call, leading incident response, debugging distributed systems under pressure