Наши требования
Expert in some differentiable array computing framework, preferably PyTorch
Expert in optimizing machine learning models for serving reliably at high throughput, with low latency
Significant systems programming experience; ex. Experience working on high-performance server systems—you’d be just as comfortable with the internals of VLLM as you would with a complex PyTorch codebase
Significant performance engineering experience; ex. Bottleneck analysis in high-scale server systems or profiling low-level systems code
Always up to date on the latest techniques for model serving optimization
Familiarity with high-performance LLM serving; ex. experience with VLLM, SGlang deployment, and internals
Experience with a public cloud platform such as GCP, AWS, or Azure
Experience deploying and scaling inference workloads in the cloud using Kubernetes, Ray, etc
You like to ship and have a track record of leading complex multi-month projects without assistance
You’re excited to learn new things and work in a multitude of roles