Program ownership: Lead planning and execution of cross-functional programs spanning data collection, annotation pipelines, alignment workflows (RLHF, DPO, Constitutional AI), safety guardrails (adversarial testing, red-teaming), and model serving. Establish scopes, goals, timelines, risks, and success metrics
Cross-functional coordination: Serve as the connective tissue between Post-Training, Safety Engineering, Trust & Safety, ML Infra, UXR, and Product. Translate model development, safety, and user experience priorities into executable roadmaps, keeping tightly coupled workstreams aligned from post-training through to production deployment
Evaluation & quality: Develop and maintain custom evaluation frameworks to track model performance and user satisfaction. Drive comprehensive quality evaluation initiatives alongside rigorous safety and toxicity baselines. Partner with UXR, researchers, and engineers to identify quality signals, incorporate human feedback, and surface actionable insights on model behavior in production
Operational excellence: Drive visibility into data pipeline health, annotation quality, training run progress, and deployment readiness. Identify bottlenecks across teams and lead efforts to improve tooling, process, and developer velocity
Strategic partnership: Partner with research, safety, product, and UXR leadership on prioritization, sequencing, and tradeoffs—balancing aggressive capability scaling with strict safety requirements, user needs, and infrastructure constraints
Process development: Build and refine the operational patterns, ontologies, and frameworks used to scale new capability development—from prompt engineering and data generation to model behavior specification and safety guidelines
Vendor & partner management: Own external partner relationships supporting these workstreams, including general and safety-focused annotation vendors, evaluation tooling providers, and data partners