5+ years of experience in ML/AI systems, with at least 2 years focused on LLMs and generative AI
Deep knowledge of the LLM ecosystem, including model architectures and fine-tuning approaches
Hands-on experience with
Running LLMs in production: deploying and operating inference workloads
LLM fine-tuning, including supervised fine-tuning (SFT/LoRA) and data preparation/curation; experience with RL-based fine-tuning is a strong plus
LLM evaluation: building task-specific benchmarks and offline/online eval pipelines, including LLM-as-a-judge setups
Inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM, Transformers)
Deploying LLM-powered applications using APIs from OpenAI, Anthropic, or open-source models
Strong Python programming skills
Excellent communication skills, with the ability to clearly explain technical concepts to diverse audiences
Must be fluent in Mandarin Chinese
Experience with inference frameworks and libraries (e.g., vLLM, SGLang, TensorRT-LLM)
Work with multimodal AI models (e.g., vision-language, speech)
Proficiency with DevOps tools (Docker, Kubernetes)
Contributions to open-source ML/AI projects