8+ years of experience in ML/AI systems, with at least 4 years focused on LLMs and generative AI
Demonstrated technical leadership: owning ambiguous, high-impact problems end to end and influencing decisions across teams and customers
Expert knowledge of the LLM ecosystem: model architectures, fine-tuning approaches, and inference internals
Deep, hands-on command of inference optimization: quantization, KV-cache management, batching, routing, etc
Hands-on experience with
Running LLMs in production at scale: deploying, operating, and debugging inference workloads down to the framework level
LLM fine-tuning, including SFT/LoRA and data preparation/curation; experience with RL-based fine-tuning
LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups
Inference frameworks and libraries (vLLM, SGLang, TensorRT-LLM), including the ability to read, modify, and contribute to their internals
Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models
Strong Python programming skills
Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences, from engineers to executives
Contributions or maintainership in major OSS inference/ML projects (vLLM, SGLang, TensorRT-LLM)
Published research, conference talks, or widely-read technical writing in the LLM/serving space
Deep work with multimodal AI models (vision-language, speech)
Proficiency with DevOps tooling (Docker, Kubernetes) and infrastructure-as-code
Experience building or owning internal tooling/automation for ML workflows at scale