This role exists to make sure the systems we’re building today can be trusted in production tomorrow, and to set the bar for what “production-ready” means on this team. The engineer in this role delivers their mandate by writing production code, shaping architecture, and engineering the systems themselves — not by absorbing operational load
YOU'LL BE RESPONSIBLE FOR
Define and operate against SLOs. Establish meaningful SLIs and SLOs with product and engineering partners, manage error budgets, and use them as real inputs to prioritization rather than dashboards no one reads
Build the observability layer. Improve metrics, logs, traces, and alerting so issues are detected early, attributed precisely, and debugged with code-level confidence. Push instrumentation upstream into the services we own
Lead incident response. Act as incident commander when needed, drive blameless postmortems, and turn findings into concrete engineering work that lands. Build the muscle in the team so this isn’t centralized in any one person
Reduce toil through engineering. Identify repetitive operational work and eliminate it with software — automation, self-healing behavior, better defaults, better tooling — rather than absorbing it as ongoing overhead
Production Hardening. Stress-test designs for partial failure, dependency degradation, traffic spikes, and adversarial inputs. Run capacity and performance work before incidents arise. Ensure resiliency primitives are tuned and working correctly
Make change safe and fast. Improve release safety through progressive delivery, feature flags, canaries, rollbacks, and tested migrations. Help the squad ship faster and with lower blast radius
Improve developer experience especially where it removes operational friction or improves change safety. Where internal tooling or platform gaps slow the team down, build or contribute the fix. Prefer leverage over heroics
Partner across disciplines. Work closely with product, platform, security, compliance, and other engineering teams. Translate reliability and risk tradeoffs into language each audience can act on
Raise the engineering bar. Mentor engineers, review hard designs and PRs, and shape technical standards across the squad. Lead through clarity and judgment, not authority
WE ARE LOOKING FOR A PERSON WHO HAS
Track record of owning services in production — not just shipping them, but being the engineer responsible for how they behave under real load and real failure
Experience defining and operating against SLOs/SLIs, and using error budgets to influence engineering and product decisions
Experience leading incident response and writing postmortems that produced durable improvements
Hands-on experience with observability tooling (metrics, structured logging, distributed tracing) and using it to diagnose nontrivial production issues
Deep system design experience: distributed services, asynchronous messaging, storage tradeoffs, API design, idempotency, consistency, backpressure, and graceful degradation
Significant industry experience building and operating production software systems in a high-ownership engineering environment
Comfort operating in modern cloud environments (e.g., AWS/GCP), containerized workloads, and CI/CD pipelines, and reasoning about their failure modes
Demonstrated technical leadership: influencing architecture across teams, mentoring strong engineers, and making the people around you better
Pragmatism. You can hold a high reliability bar while still helping a fast-moving squad ship