You have 5+ years of people-leadership experience and 8+ years total in technical support/operations, including running a 24/7 support function at scale in a cloud operations environment
You have a strong background in Linux, containerization technologies, and Kubernetes, and you understand virtualization and cloud computing concepts. Experience at a hyperscaler or cloud infrastructure provider is a strong plus
You lead with empathy and aren't afraid to get your hands dirty. You do the work alongside your direct reports and model the standard you set
You're energized by leadership excellence and talent development: diligent performance management, coaching, and growing each engineer according to their individual needs
You've built enablement and quality programs that scaled across a function or multiple teams — and can show the measurable improvement they drove
You've designed the operating model for a support org, including coverage/staffing model, escalation boundaries, SLOs, and evolved it as the org scaled
You're a calm, clear communicator with customers and executives during critical incidents, and you resolve conflicts effectively across teams and organizational boundaries
You think in systems and multi-quarter plans. You've defined KPI/SLO frameworks, reported to senior leadership, and owned capacity and growth planning for a function, not just a single team
You've led a globally distributed team across time zones
Leadership & Communication: proven ability to lead through senior talent and set direction for a function, with executive-level communication skills
Strategic & Operational Planning: you build operating systems, plans, and metrics that scale a function beyond what any one person can hold
Problem-Solving & Adaptability: robust problem-solving skills and adaptability at organizational scale, in a fast-paced, hyper-growth environment
Program Management: experience with program-management tools and methodologies
Experience supporting AI/ML, HPC, or GPU-accelerated workloads at scale
Hands-on Kubernetes operations experience (CKA certification a plus)
Familiarity with Slurm/SUNK, RDMA networking, distributed storage, and observability tooling such as Grafana
You have some experience with infrastructure as it relates to Data Center Operations
Wondering if you're a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we've found compatible with our team. If some of this describes you, we'd love to talk
You're an expert in what it takes to be an excellent leader and foster an environment in which people are excited and inspired to participate
You love to dive into problems, test for solutions, and enjoy engaging with customers
You're excited and curious about AI
Why CoreWeave?
At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values
Be Curious at Your Core
Act Like an Owner
Empower Employees
Deliver Best-in-Class Client Experiences
Achieve More Together
We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!