Наши требования
Experience: 5+ years engineering experience, with a strong track record in inference acceleration and model deployment at scale
Inference Mastery: Proven expertise in inference optimization, including quantization, attention acceleration, and deep learning compiler stacks
GPU & Parallelism: Deep knowledge of GPU programming (CUDA, NCCL) and experience with SP, TP, PP, and other forms of parallelism for distributed inference
AI Domain Knowledge: Familiarity with video generation (videogen) models and large language models (LLMs)
Collaboration: Strong cross-discipline communication skills; able to drive shared goals across research and engineering functions
Ownership Mindset: Self-driven, solutions-oriented, and capable of managing ambiguity in a fast-paced startup environment
Bonus: Experience in enhancing training efficiency, stability, or resource optimization for large models
Experience with high-throughput video or real-time streaming model deployment
Familiarity with distributed training and optimization toolkits
Contributions to open source projects in AI infrastructure or deep learning compilers
Startup or rapid prototyping experience