Define meaningful SLIs, SLOs, and reliability targets for the platform
Collaborate with the software engineering teams to define and achieve the best practices for software observability, SLIs, SLOs and reliability
Strengthen production readiness by improving service ownership, observability, alerting, runbooks, scaling assumptions, rollback paths, and failure-mode preparedness
Improve the reliability, scalability, and performance of cloud, Kubernetes, and shared infrastructure
Build actionable observability using metrics, logs, traces, and golden signals, with tools such as Datadog, Prometheus, and Grafana
Implement operational and security best practices through guidelines, policies and automation
Reduce alert noise and improve signal quality so teams can detect, understand, and resolve issues quickly
Automate repetitive operational work using Python or other languages
Implement self-service Internal Developer Platform features via APIs and Kubernetes operators
Improve deployment safety, rollbackability, and release observability
Improve reliability of critical stateful systems such as databases, caches, queues, and streaming platforms
Participate in on-call, troubleshoot, and coordinate incident response, and facilitate blameless post-incident reviews that turn into concrete improvements
Execute disaster recovery drills and analyse cloud/platform usage to identify cost and resource-efficiency gains without compromising reliability