5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI
Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches
Hands-on experience with
Running LLMs in production: deploying and operating inference workloads
LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation; experience with RL-based fine-tuning is a strong plus
LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups
Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers)
Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models
Strong Python programming skills
Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences
Experience with inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM)
Work with multimodal AI models (e.g., vision-language, speech)
Proficiency with DevOps tools (Docker, Kubernetes)
Contributions to open-source ML/AI projects