3–5+ years in Product Management, ML infrastructure/MLOps, distributed systems, or cloud platform engineering
Strong technical depth in distributed systems, cloud infrastructure, or ML platforms
Hands-on familiarity with large-scale ML training and orchestration tools (e.g., Slurm, Kubernetes, Ray)
Track record of shipping technically complex products with multiple engineering teams
Strong communication and stakeholder management across engineering, research, and customers
Experience with product analytics, data-informed prioritization, and experimentation
High ownership, high learning velocity, and comfort operating in fast-moving AI infrastructure environments
Experience with GPU platforms and HPC primitives: InfiniBand/RDMA, topology-aware scheduling, high-throughput storage
Practical understanding of modern ML training stacks: PyTorch, DeepSpeed, FSDP/ZeRO, NCCL
Familiarity with efficiency and reliability metrics: Goodput, MFU, failure modes, preemption handling, health checks
Exposure to large-scale LLM training/inference systems
Experience in observability, performance tuning, or SRE/reliability engineering
Customer-facing technical experience (solutioning, support, architecture advisory)