Frontier AI companies are increasingly bottlenecked on expert judgment — capturing it reliably, validating it at scale, and turning it into durable model behavior. This role sits at the center of that problem
You'll build the ML systems that power Mercor's Frontier Data Products: the infrastructure that scores, validates, and improves complex work products where correctness is rarely binary and labels are often noisy, delayed, or disputed. A single job can stay live for days, interleaving model inference, automated checks, expert review, disagreement resolution, and feedback loops. Your work determines how models reason over ambiguous inputs, when they should defer to humans, how quality is measured, and how feedback compounds into better systems over time
This is applied ML product engineering under real production constraints — incomplete ground truth, shifting requirements, latency and cost tradeoffs, and workflows where a silent model failure corrupts the final output. It is not an offline benchmarks role
Build ML systems that score, validate, and improve complex work products where correctness is nuanced and labels are imperfect
Design evaluation frameworks for ambiguous tasks where ground truth is partial, delayed, or disputed
Build feedback loops that turn review, disagreement, correction, and adjudication into measurable model and system improvements
Own production ML behavior end-to-end: precision/recall tradeoffs, regression detection, drift, latency, cost, and explainability
Improve model quality using the right tool for the job — prompting, fine-tuning, retrieval, active learning, heuristics, and error analysis
Partner with backend engineers to integrate inference into durable, long-running workflows without sacrificing debuggability or human oversight
Moving fast on a young, high-ownership codebase where your decisions have long-term architectural weight
Operating across models, data, backend systems, and product surfaces — context switching is the default, not the exception
Debugging production ML failures in live, long-running workflows where silent errors matter
Working closely with backend engineers on a stack of Python, Temporal, Postgres, AWS, and LiteLLM
Balancing automation confidence with human review — knowing when to defer is as important as knowing when to ship