Чем предстоит заниматься
This is a high-impact role where you'll balance reliability with velocity, knowing when to move fast and when to prioritize stability. You'll lead incident response, drive systemic improvements, and help shape how Gamma scales to serve its next 100 million users
Our team has a strong in-office culture and works in person 4–5 days per week in San Francisco. We love working together to stay creative and connected, with flexibility to work from home when focus matters most
Own the reliability, availability, and performance of Gamma's production systems across our AWS infrastructure
Build observability infrastructure from the ground up: metrics, logging, tracing, and alerting that give the team genuine visibility into system health before users feel the impact
Design and ship automation that reduces toil, makes deployments safer, and gets us back on our feet faster when things go wrong
Lead incident response and blameless post-mortems, then follow through on the systemic fixes that keep the same issues from coming back
Partner with engineering teams on architecture reviews, SLO and SLI design, and reliability best practices that scale with the product
Manage and optimize our compute, networking, databases, and managed services