8+ years of experience building and operating large-scale distributed systems or cloud services
3+ years of engineering management experience, including managing senior and staff-level engineers
Track record owning a production service end to end including availability, latency, capacity, and cost
Deep expertise in Kubernetes at production scale, including orchestration, scheduling, and service design
Strong understanding of distributed systems, networked services, and performance under load; able to lead architecture and code review without owning the critical path
Experience defining and defending SLIs/SLOs and error budgets, and building the operational practice behind them
Experience with multi-tenant platforms: isolation, quota, noisy-neighbor mitigation, and fair sharing
Excellent communication and stakeholder management; able to create clarity in ambiguous or fast-scaling environment
Demonstrated track record of growing and promoting engineers, including into senior and staff-level roles
Experience running performance calibration and making accurate leveling decisions
Experience managing underperformance directly and fairly, as well as retaining top performers
Experience building and owning a hiring bar and interview process for senior engineering roles
Experience operating a customer-facing cloud service or developer platform, including API and control-plane design
Familiarity with inference-serving stacks such as vLLM, Triton, TensorRT-LLM, Ray Serve, or TorchServe
Experience operating multi-region fleets at a cloud provider or hyperscaler
Familiarity with observability stacks (Prometheus, Grafana, OpenTelemetry)
Experience scaling teams, tooling, and processes in a high-growth environment
Wondering if you’re a good fit? We believe in investing in our people, and value candidates who can bring their own diversified experiences to our teams – even if you aren't a 100% skill or experience match. Here are a few qualities we’ve found compatible with our team. If some of this describes you, we’d love to talk
You love to build frictionless products for developers
You’re curious about AI and MLOps tooling
You’re an expert in building inference systems that scale for production workloads
Why Us?
At CoreWeave, we work hard, have fun, and move fast! We’re in an exciting stage of hyper-growth that you will not want to miss out on. We’re not afraid of a little chaos, and we’re constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values
Be Curious at Your Core
Act Like an Owner
Empower Employees
Deliver Best-in-Class Client Experiences
Achieve More Together
We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and enables the development of innovative solutions to complex problems. As we get set for takeoff, the organization's growth opportunities are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!
Experience operating low-latency, high-throughput request-serving systems against strict P95/P99 targets
Working knowledge of autoscaling, load balancing, admission control, and traffic management at scale
Experience with capacity planning and demand forecasting for constrained or expensive resources
Familiarity with GPU-backed workloads and the scheduling and utilization trade-offs they introduce
Progressive delivery practices canarying, feature gating, safe rollback for always on services