Наши требования
MS or PhD in Computer Science, Mathematics, or a related field
5+ years of relevant, hands-on industry experience
Proficiency in C++
Experience writing or improving kernels (Triton, CuTeDSL, TileLang, CUDA, CUTLASS, ThunderKittens) to resolve low-level bottlenecks
Proven success deploying performant inference at scale using open-source or custom inference engines, routers, etc
Direct experience scaling models via FSDP, Tensor Parallelism, or related sharding techniques on multi-node GPU clusters
Experience designing reinforcement learning systems for high-throughput training and asynchronous data sampling
BS in Computer Science or a related technical field, or equivalent industry experience
2+ years of relevant, hands-on industry experience
Proficiency in Python
Experience building or maintaining components within ML frameworks (e.g., PyTorch, JAX, or TensorFlow)
Proficiency in either
Understanding of distributed training concepts and collective communication primitives (e.g., NCCL)