Bachelor's degree or higher in Computer Science or related field
Proven track record owning production-grade real-time, large-scale systems where tail latency (p99) matters
Proficient coding abilities in one or more popular programming or scripting languages; Python proficiency is a plus
Good taste in product, particularly developer-oriented tools
Interest in ML/AI infrastructure and willingness to learn
Strong collaboration and communication skills
Comfortable using AI coding assistants (e.g., Claude Code, Codex, Cursor) as a daily productivity multiplier — as an AI-native company, we see this as a must-have skill
Experience implementing pipeline-level model runtime optimizations such as dynamic batching, async scheduling, or decode-side throughput improvements
Experience building developer platforms: SDKs, CLIs, APIs, and self-serve workflows for ML or infrastructure products
Experience with containerization and orchestration technologies (Docker, Kubernetes), service meshes, or distributed scheduling
Familiarity with speech/audio ML models (STT, TTS, speech-to-speech)
Familiarity with model-serving runtimes (vLLM, TensorRT, ONNX)
Familiarity with systems-level performance profiling across host-device boundaries (e.g. PyTorch Profiler), diagnosing GPU utilization issues
Exposure to customer-facing engineering: pre-sales prototyping, technical discovery, or working directly with customers to ship solutions