L

Research Scientist / Engineer – Reinforcement Learning Infrastructure

Luma Ai
≈ 1,3 млн–2 млн ₽ · 14,6 тыс.–21,9 тыс. €
16 667 – 25 000 $
🌍 Удалённо
Senior
Полная занятость
🌐 Глобал
Продуктовая
Tech
Описание вакансии
Обязанности
Build high throughput rollout generation
Build reward model serving pipelines
Collaborate with researchers to operationalize RL training ideas
Create LLM as judge evaluation pipelines
Design curriculum and task sampling
Design distributed reinforcement learning post training systems
Develop evaluation monitoring and debugging tooling
Develop reward infrastructure and verifiable rewards
Implement defenses against reward hacking
Implement reinforcement learning environments for multi step tasks
Implement sequence packing for long trajectories
Improve training efficiency and stability
Integrate inference engines into training loop
Reuse KV cache across rollouts
Schedule heterogeneous training and inference workloads

Технологии: Asynchronous training, Containerization, Curriculum learning, Distributed Systems, Distributed Training, FSDP, GPU Cluster, Human Feedback, KV cache, Kubernetes, LLM Evaluation, Language Models, Large Language Models, Learning from Human Feedback, MPI, Mixture of Experts, Model Synchronization, NCCL, Off Policy, Off Policy Learning, Pipeline parallelism, Policy Optimization, Policy learning, Proximal Policy Optimization, PyTorch, Ray, Reinforcement Learning, Reinforcement Learning from Human Feedback, Reward Hacking, Reward Modeling, SGLang, Sandboxing, Sequence Packing, Tensor Parallelism, Tool use, VLLM

AI-помощник
ИсточникСкрыто
Опубликовано3 дн назад
Выше рынка на 384%
вакансия1 307 000
в среднем по рынку270 000
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение — прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).