WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Technical Program Manager - Cluster Orchestration & Applied Training
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Technical Program Manager - Cluster Orchestration & Applied Training

CoreWeave·Bellevue, WA, Sunnyvale, CA, New York, NY, Livingston, NJ·6 мая

Technical Program Manager - Cluster Orchestration & Applied Training

🏢 ОфисMiddleПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Чем предстоит заниматься

Drive end-to-end program execution for cluster orchestration initiatives spanning workload scheduling, self-service provisioning, upgrade and migration flows, and platform integrations
Lead cross-functional programs that improve how AI training, evaluation, RL, and mixed workloads run across CoreWeave clusters
Partner with engineering and product leaders to define roadmap priorities and deliver measurable improvements in utilization, reliability, scalability, observability, and user experience
Drive delivery for applied training initiatives across pre-training, fine-tuning, reinforcement learning, sandbox environments, and evaluation systems
Coordinate dependencies across platform engineering, infrastructure, product, customer-facing teams, and ecosystem partners to ensure successful launches and clear operational ownership
Build program mechanisms for release readiness, rollout planning, risk management, stakeholder communication, and post-launch review
Establish success metrics, dashboards, and operating cadences to improve cluster efficiency, workload startup performance, time-to-research, and adoption of new platform capabilities
Create clarity across ambiguous technical programs by aligning stakeholders, surfacing tradeoffs early, and driving decisions to resolution
Bachelor’s degree in Computer Science, Engineering, or a related technical field, or equivalent practical experience
8+ years of technical program management experience in cloud infrastructure, distributed systems, or AI/ML platforms
Experience leading large-scale cross-functional programs involving scheduling systems, cluster infrastructure, or ML platform capabilities
Strong technical fluency in Kubernetes, Slurm or comparable schedulers, distributed systems, and AI training workflows
Demonstrated ability to define program metrics and deliver measurable outcomes in performance, reliability, scale, or operational maturity
Excellent communication skills, with experience influencing engineering, product, and executive stakeholders
Experience with orchestration and scheduling technologies such as Kubernetes, Slurm, Kueue, Ray, or similar systems
Familiarity with modern AI training and evaluation workflows, including pre-training, supervised fine-tuning, reinforcement learning, and experiment or sandbox environments
Understanding of GPU infrastructure, cluster capacity planning, multi-tenant execution, and distributed training tradeoffs
Experience building launch processes, release governance, dependency management, and operational review mechanisms in fast-scaling environments
Familiarity with AI developer and research tooling such as W&B, SkyPilot, or adjacent ecosystem platforms
Be Curious at Your Core
Act Like an Owner
Empower Employees
Deliver Best-in-Class Client Experiences
Achieve More Together
Medical, dental, and vision insurance - 100% paid for by CoreWeave
Company-paid Life Insurance
Voluntary supplemental life insurance
Short and long-term disability insurance
Flexible Spending Account
Health Savings Account
Tuition Reimbursement
Ability to Participate in Employee Stock Purchase Program (ESPP)
Mental Wellness Benefits through Spring Health
Family-Forming support provided by Carrot
Paid Parental Leave
Flexible, full-service childcare support with Kinside
401(k) with a generous employer match
Flexible PTO
Catered lunch each day in our office and data center locations
A casual work environment
A work culture focused on innovative disruption

Дополнительно

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025
What You’ll Do
CoreWeave is seeking a Technical Program Manager to lead complex, cross-functional programs across Cluster Orchestration and Applied Training within our AI/ML Platform Services organization
Cluster Orchestration is the platform layer that makes sure large AI workloads are scheduled, launched, and managed reliably across CoreWeave’s clusters. Applied Training is the layer on top of that infrastructure that helps researchers and customers use it for pre-training, fine-tuning, reinforcement learning, evaluations, and sandboxed experimentation
About the role
In this role, you will partner with engineering, product, infrastructure, and research-adjacent teams to improve both how workloads run on the cluster and how users interact with the training platform built on top of it. That includes driving programs across orchestration systems such as Slurm-on-Kubernetes (SUNK), Kueue, and workflow integrations, while also helping scale the environments, tooling, and operational mechanisms that make training and evaluation workflows easier to use
This is a highly cross-functional role for a TPM who combines strong technical depth, excellent execution instincts, and the ability to bring structure and clarity to fast-moving infrastructure and AI platform initiatives
C
CoreWeave
Bellevue, WA, Sunnyvale, CA, New York, NY, Livingston, NJ

ГрейдMiddle
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано6 мая
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).