PhD in computer science, machine learning, ML systems, computer systems, computer architecture, electrical engineering, applied math, or a closely related field
Strong publication record or equivalent research artifacts in ML, ML systems, efficient inference, model compression, quantization, distillation, serving systems, or related areas
Strong hands-on coding ability in Python and PyTorch; ability to move from idea to experiment to prototype quickly
Deep understanding of LLMs, VLMs, transformer inference, decoding algorithms, model compression, quantization, and production-serving tradeoffs
Strong experimental design skills, including ablations, baselines, metrics, statistical reasoning, and failure analysis
Excellent written and verbal communication
First-author publications in NeurIPS, ICML, ICLR, MLSys, ACL, EMNLP, ASPLOS, OSDI, SOSP, ISCA, HPCA, or comparable venues
Experience deploying ML models or inference optimizations in production
Experience with vLLM, SGLang, TensorRT-LLM, NVIDIA Dynamo, FlashAttention, FlashInfer, Triton, CUDA, or PyTorch internals
Experience with post-training, SFT, DPO, RLHF, RLAIF, preference optimization, or synthetic data generation when connected to inference quality or efficiency
Open-source research artifacts, widely used benchmarks, high-quality technical blogs, or invited talks in efficient AI systems