PhD in Computer Science, Machine Learning, AI, or a related field, or equivalent industry experience
Significant experience leading technical projects or teams in machine learning, AI research, or large-scale distributed systems. Experience scaling and mentoring high-performing research and engineering teams
Deep understanding of modern machine learning techniques, including transformers, reinforcement learning, alignment methods, and large language models
Strong track record of delivering impactful research or applied ML systems in production environments
Expertise in designing, building, and maintaining production-quality ML systems and infrastructure
Experience training, serving, debugging, and optimizing large-scale models on GPU-based systems
Experience leading teams working on large language model training, mid-training, or post-training
Experience with product experimentation, online evaluation, and A/B testing frameworks
Strong software engineering skills with the ability to write clean, maintainable, and scalable code
Excellent communication skills and the ability to influence technical direction across teams. Lead complex, cross-functional initiatives across data, training infrastructure, evaluation, and model serving
Hands-on experience working directly with open-source models like Mistral and Qwen, particularly adapting them via mid- and post-training for specific personas, creative writing, or role-playing applications
Familiarity with cloud-native ML infrastructure, including Kubernetes, Docker, and modern orchestration platforms
Publications in leading machine learning conferences or demonstrated contributions to the broader AI community