8+ years of software engineering experience with strong fundamentals in distributed systems, system design, data structures, and algorithms
Strong Python skills and a track record of shipping production software; comfort in at least one other part of the stack (TypeScript/React, Go, Rust, or similar)
Deep experience with containerization and sandboxed execution, including Docker, VMs, gVisor/Firecracker, Kubernetes, or equivalent
Experience building or operating high-throughput backend systems: orchestration, job scheduling, queuing, and large-scale data pipelines
Hands-on experience building with LLMs including agent loops, tool calling, MCP, or eval harnesses, and enough intuition about model behavior to reason about what a training signal actually teaches
Demonstrated ability to own ambiguous, undefined problems end to end and drive them to a shipped system
Excellent written and verbal communication; ability to align engineers, researchers, and non-engineering partners on a technical direction
Direct experience building RL environments, agentic benchmarks, or eval harnesses (SWE-bench-style task suites, terminal or browser environments, tool-use benchmarks, or in-house equivalents)
Familiarity with post-training methods: RLHF, RLAIF, RLVR, GRPO/PPO-family algorithms, rejection sampling, reward modeling, and the practical failure modes of each
Experience designing verifiable reward signals, and firsthand experience with reward hacking and how to defend against it
Experience with RL training or serving stacks (verl, TRL, Ray, vLLM, SGLang, or similar)
Experience with high-scale sandbox or code-execution infrastructure, remote development environments, or CI systems
Experience with cloud-native infrastructure across AWS/GCP/Azure, Infrastructure as Code, and CI/CD
Strong observability instincts: tracing, structured logging, and metrics for systems whose failure modes are statistical rather than binary
Experience building internal tools that non-engineers rely on daily, especially data-dense review and annotation interfaces
Experience in a research-adjacent engineering role, translating research goals into production systems
Experience working directly with sophisticated external technical customers
Prior technical leadership at staff level or above in a fast-moving, ambiguous environment
PLEASE NOTE: Our policy requires a 90-day waiting period before reconsidering candidates for the same role. This allows us to ensure a fair and thorough evaluation of all applicants