AI and Machine Learning Infrastructure: You understand how training, fine-tuning, inference, checkpointing, and GPU-intensive workloads interact with orchestration systems
Strong Technical Fluency: Significant experience with Kubernetes, cloud infrastructure, distributed systems, workload orchestration, or related technologies
Technical Product Management Experience: Five or more years of experience owning complex technical products or infrastructure initiatives from definition through delivery and iteration
Systems Thinking and Attention to Detail: The ability to reason through APIs, lifecycle states, failure modes, upgrades, operational workflows, and dependencies across distributed systems
Strong Product Judgment: The ability to balance customer value, reliability, engineering cost, operational complexity, and long-term maintainability when making product decisions
Ownership and Execution: A track record of independently driving ambiguous, cross-functional initiatives to clear decisions and high-quality outcomes
Clear Communication and Collaboration: The ability to work closely with engineers, document decisions precisely, and create alignment across technical and nontechnical teams
High-Performance Computing: You have experience with HPC environments, Slurm or similar workload managers, specialized networking, or large accelerator clusters
Infrastructure Operations: You have operated or supported production infrastructure and understand the impact of product decisions on reliability, upgrades, debugging, and incident response
Builder Orientation: You enjoy using infrastructure products directly, examining APIs, and experimenting with systems to develop grounded product opinions