Наши требования
5+ years of research experience in large-scale pre-training/mid-training of multimodal foundation models (LLMs, VLMs, Audio LMs, or similar), ideally at the staff or lead scientist level
Track record as a first author on major publications in top conferences or journals (e.g., NeurIPS, ICML, ICLR)
Extensive hands-on experience with large-scale multimodal model design, training, and deployment
Deep understanding and implementation experience with generative architectures (diffusion, autoregressive, cross-modal, etc.)
Expertise in high-throughput, scalable dataset curation and model pipeline optimization for multimodal applications
Strong programming and prototyping skills (Python, PyTorch, TensorFlow, etc.) and experience deploying research into production systems
Excellent communication and collaboration skills, and a passion for building creative enabling technology