We're looking for a Data Scientist to take ownership of the datasets and evaluation workflows that underpin our AI Safety work. You'll turn real-world product data into reliable, well-structured datasets for training and evaluating models, and help build the processes and infrastructure to continuously assess how those models perform in production
This is a hands-on role at the intersection of data science, ML, AI safety, and policy. You'll work closely with safety researchers, engineers, and policy specialists to translate complex safety requirements into practical data and evaluation systems - starting hands-on with collection, analysis, and labelling, then building the methodologies, contributor networks, and quality standards that let this work scale
Own safety datasets end-to-end: collection, cleaning, labelling, quality control, versioning, and readiness for training and evaluation
Translate safety policy into clear, consistent labelling and evaluation criteria, working closely with policy specialists
Design and manage labelling processes, including sourcing, onboarding, and overseeing external contributors to a high quality bar
Build evaluation workflows for models in production, using real-world data to track performance and surface issues
Develop lightweight Python/SQL pipelines and tooling to make data work faster and reproducible, partnering with ML engineers on what "training-ready" looks like