Наши требования
Experience building SRE practices from scratch
Strong hands-on SRE engineering experience
Good understanding of SLI, SLO, error budget, burn rate, and SLO-based alerting
Experience with incident management, on-call, escalation, postmortems, and runbooks
Experience with monitoring and observability tools such as Prometheus, Grafana, Splunk, Datadog, or similar
Understanding of OpenTelemetry and distributed tracing concepts
Experience with cloud platforms, preferably Azure
Experience with Kubernetes, Terraform, and CI/CD pipelines
Strong troubleshooting skills across application, infrastructure, and platform layers
Ability to communicate clearly with technical teams and client stakeholders
Strong spoken and written English skills (Intermediate level and more)