Наши требования
Deep ML systems background (e.g., training compilers, runtime optimization, custom kernels)
Experience operating close to hardware (GPU/TPU performance tuning)
Background in robotics, multimodal models, or large-scale foundation models
Experience designing abstractions that balance researcher flexibility with system reliability
Pursuant to the San Francisco Fair Chance Ordinance, we will consider for employment qualified applicants with arrest and conviction records
Strong software engineering fundamentals and experience building ML training infrastructure or internal platforms
Hands-on large-scale training experience in JAX (preferred), PyTorch
Familiarity with distributed training, multi-host setups, data loaders, and evaluation pipelines
Experience managing training workloads on cloud platforms (e.g., SLURM, Kubernetes, GCP TPU/GKE, AWS)
Ability to debug and optimize performance bottlenecks across the training stack
Strong cross-functional communication and ownership mindset