As the Senior Site Reliability Engineer, you will lead Branch’s effort to achieve greater reliability, performance, scalability, capacity and observability of our platform through automation and tooling. You will also participate in and improve the software development and deployment life cycles as well as develop and implement technical best practices
Partner with Developers to produce high-performing and robust services through rigorous testing and release procedures
Design infrastructure, monitoring, processes, and standards for systems and applications
Support services through design, development, load testing, and launch phases
Develop, measure, and monitor key performance and service level indicators including availability, latency, and overall system health
Define and establish SLIs, SLOs, and error budgets with service owners, and drive adoption across platform teams
Profile and optimize platform performance, resilience, and efficiency, including latency, throughput, and capacity planning under load
Participate in incident response and root cause analysis
Remediate tasks and develop preventative and automated measures to meet SLAs/SLOs/SLIs
Manage monitoring services utilized by applications