As the Engineering Manager for Baseten's Cloud Platform team, you will directly manage a team of cloud platform engineers responsible for building the systems and processes that keep our infrastructure scalable, reliable, and efficient — from automated deployments and monitoring to performance optimization and incident response
You are a people-first leader with a strong cloud infrastructure background. You set a high bar for reliability and operational excellence, engage credibly in technical discussions and code reviews, and know how to build a culture of ownership and accountability. You'll spend most of your time close to the work: unblocking your team, shaping technical direction on day-to-day decisions, and developing your engineers. At Baseten, we work closely with our users to understand their struggles operationalizing ML — you'll keep your team connected to that mission and translate user learnings into better infrastructure
Recruit, hire, and grow a high-performing team of cloud platform engineers; provide ongoing coaching, feedback, and career development through regular 1:1s
Set clear performance expectations, hold a high bar, and create an environment where engineers do their best work
Foster a culture of ownership, accountability, and continuous improvement
Drive day-to-day technical decisions through design reviews, code reviews, and architectural discussions; translate the infrastructure roadmap into clear team priorities and milestones, and hold execution against them
Establish standards and best practices for reliability, performance, and operational excellence; ensure the team owns projects end-to-end, from specification through production
Exercise and encourage good judgment on tooling and architectural tradeoffs, with a bias against unnecessary complexity
Partner with product and engineering teams to align infrastructure work with business priorities; represent your team's progress, capacity, and technical tradeoffs clearly to leadership
Serve as the escalation point for major incidents; drive resolution with urgency, ensure the team learns systematically, and stay connected to users so those learnings feed back into infrastructure improvements