Strong machine learning and engineering background
Experience with Large Language Models (LLM), including
Understanding of how LLMs learn
Data ablations and scaling laws
Post-training techniques
Training reasoning and agentic models
Experience with implementing cost-efficient, complex pipelines to generate synthetical datasets at scale optimizing for data quality, correctness, diversity, etc
Experience with evals tracking model capabilities (general knowledge, reasoning, math, coding, long-context, etc)
Experience in building trillion-scale pretraining datasets, and familiarity with concepts like data curation, deduplication, data mixing, tokenization, curriculum, impact of data repetition, etc
Excellent programming skills in Python
Strong prompt engineering skills
Experience working with large-scale GPU clusters and distributed data pipelines
Strong obsession with data quality
Research experience
Author of scientific papers on any of the topics: applied deep learning, LLMs, source code generation, etc
Formal machine learning, mathematics, or computer science background
Can freely discuss the latest papers and descend to fine details
Is reasonably opinionated