As a Site Reliability Engineer at Baseten, you'll define and codify the gold standards of day 2 operations for our ML infrastructure platform. You'll envision and build robust systems, processes, automations, and observability tooling that keep our platform reliable at scale — and that empower the broader organization to operate confidently
You'll work closely with engineering, forward-deployed and product teams: learning from recurring failure patterns, turning tribal knowledge into automated mitigations, and raising the operational floor for the entire company
EXAMPLE INITIATIVES
You'll work on projects like these as part of the SRE team
Improve Baseten SRE Practices, by instrumenting SLOs and SLIs, improving alerting and observability for all services
Building AI-assisted tooling for incident triage and response
Own the reliability of Baseten's multi-cloud Kubernetes infrastructure, including incident response, post-mortems, and remediation tracking
Build and maintain observability infrastructure — metrics, logging, dashboards, and alerting — as code
Author, validate, and improve runbooks for recurring failure patterns, ensuring they're structured for low-context, safe execution
Identify high-frequency failure patterns and convert them into automated mitigations or self-healing automations
Diagnose and resolve runtime issues related to latency, memory behavior, GPU utilization, concurrency, and model lifecycle management
Define and instrument SLOs and SLIs across customer workloads and internal services
Navigate ambiguity, make principled tradeoffs, and avoid unnecessary complexity in the systems you build and the processes you define