Наши требования
SRE/production support experience in large distributed systems, preferably in banking or financial services
Deep understanding of SLI, SLO, error budget, burn rate, and SLO based alerting
Experience with incident management: on call, escalations, postmortems, and runbooks
Experience with observability tools such as Prometheus, Grafana, Splunk, Datadog, or similar
Familiarity with OpenTelemetry and distributed tracing concepts
Experience with cloud platforms (preferably Azure)
Strong troubleshooting skills across application, infrastructure, and platform layers
Remote work
Full-time (8 hours/day)
Attractive USD compensation
Paid vacation, holidays