Bachelor’s degree in Computer Science, Engineering, Statistics, or a related field. Master’s or higher preferred but not a requirement
6+ years of experience shipping software in production, including AI/LLM features
Fluency in at least one of Python, TypeScript, or Go, and willingness to work across all three
A measurement-first, systems-thinking instinct: you optimize for user outcomes over isolated metrics, and you can design an eval that is not fooling you (sampling, ground-truth quality, leakage, noise)
Comfort debugging complex, unpredictable, real-world failures
Strong communication skills: you can make a quality or cost result legible to engineers, product, and leadership, and collaborate effectively in a team environment
Deep hands-on experience with agentic coding tools and real intuition for model strengths, failure modes, and prompting limits
Prior work on eval harnesses, LLM observability, or safety/guardrails in production
Background in data engineering, data modeling, analytics, retrieval/RAG, or semantic layers, which is highly relevant for data-centric coding agents
Experience working with large-scale datasets or production system logs