Чем предстоит заниматься
Own release management end-to-end, ensuring on-time, high-quality product releases through coordination across all teams in a fast-paced environment
Introduce and enforce change safety standards, such as risk assessments, rollback procedures, feature flag, and bug bashes to reduce regressions and customer impact
Lead horizontal reliability initiatives focused on improving test coverage, observability, and incident response readiness
Define, measure, and report on reliability metrics (e.g. change failure rate, MTTR, SLI), and drive accountability for sustained improvement
Identify systemic gaps in release processes, testing, monitoring, and incident response; convert findings into structured improvement plans with clear owners and timelines
Drive rapid triage and resolution of customer-reported issues in partnership with Product and User Operations, ensuring timely followup and continuous improvements
Own and improve the incident management lifecycle: Facilitate rigorous post-incident reviews, ensuring root causes are identified and corrective and preventative actions are tracked to completion
Oversee vendor reliability and SLA compliance, including performance monitoring, incident escalation, and periodic business reviews