The ideal candidate combines strong product judgment, technical fluency, operational rigor, and customer-facing experience, with a passion for turning emerging model capabilities into a credible measure of what AI can actually do for the economy
Own the roadmap and portfolio strategy: Set and influence priorities across new benchmark development, leaderboard launches, infrastructure investment, and expansion into new evaluation categories. Decide which domains earn a slot, when a benchmark has saturated, and what replaces it. Identify opportunity areas for strategic partnership
Run the intake for new benchmarks: Evaluate and prioritize proposals from research, customers, and GTM against real demand and company strategy
Enforce eval integrity: Own contamination policy, holdout strategy, versioning, auditability, and release cadence. Publish methodology clearly enough that a skeptical researcher can reconstruct our results
Build end-to-end eval pipeline: Own the path from eval run to published result, including harness execution, grading, model onboarding, hyperparameter scaffolds, cost and latency reporting, and the leaderboard surface itself
Work directly with labs and customers: Understand evaluation needs, present and defend results, and turn what you hear into the next benchmark. Support GTM on launches, partnerships, and thought leadership to influence customer model development strategies
Close the loop to the business: Connect leaderboard demand signals to loss analysis investments and dataset production. Track adoption, usage, and downstream revenue to continuously justify leaderboard ROI
Get in the weeds: Write specs and PRDs, but also read trajectories, spot-check failures, and make small PRs to unblock yourself and continuously improve the system