WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Senior Site Reliability Engineer (SRE)
ВердиктОписаниеИнструментыКомпанияПохожие
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Senior Site Reliability Engineer (SRE)

EPAM·26 авг.

Senior Site Reliability Engineer (SRE)

🌍 УдалённоSeniorПолная занятостьАутсорс
Зарплата не указана
44
Есть о чём спросить
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Чем предстоит заниматься

Own the observability charter for the platform: build monitoring, alerting, synthetic checks, dashboards, and runbooks
Define meaningful SLIs/SLOs and reduce alert noise to improve signal quality
Design and optimize release pipelines with progressive delivery, health gates, and automated rollback mechanisms
Automate operational toil through scripting and infrastructure-as-code
Accelerate SRE maturity by applying AIOps capabilities to improve detection, diagnosis, and reduce manual effort
Lead incident response practices including on-call readiness and blameless post-mortems
Collaborate with DevOps, Cloud teams, product engineering teams, and Tech Leads to drive reliability improvements

Наши требования

3+ years of experience in Site Reliability Engineering, DevOps, or platform/production engineering supporting customer-facing systems
Expertise in observability tools such as Grafana, Prometheus, and log/trace aggregation (Loki, Tempo, OpenTelemetry) covering metrics, logs, traces, and events
Knowledge of SRE fundamentals: SLIs/SLOs, error budgets, golden signals, alert tuning and noise reduction, and blameless post-incident reviews
Experience operating workloads on Kubernetes (ideally EKS) and AWS, with the ability to debug issues across application, container, and infrastructure layers
Proficiency in Python, Bash, or Go, with exposure to infrastructure-as-code (Terraform) and CI/CD pipelines
Background in incident management including triage, escalation, communication, post-mortems, and on-call processes/rotations
A proactive ownership mindset with the ability to identify problems from telemetry before they're reported and follow through with engineering teams
Strong communication skills to turn noisy signals into crisp findings, runbooks, and recommendations
English Level: B2+ (Upper-Intermediate) or higher
Experience setting up synthetic monitoring (API and browser checks) to validate critical user journeys and catch failures proactively
Skills in performance engineering, including load/stress testing (k6, JMeter, Locust), capacity planning, and profiling latency, throughput, and resource bottlenecks
Familiarity with leveraging AIOps capabilities to advance SRE maturity and drive innovation
Experience with Datadog or similar enterprise observability platforms
Background in evangelizing best practices and setting standards across engineering teams
Exposure to programmatic advertising or adtech platforms

Дополнительно

We are looking for a Senior Site Reliability Engineer
To work hands-on with a Grafana-based observability stack and AWS/Kubernetes (EKS), defining SLIs/SLOs, reducing alert noise, building actionable dashboards, strengthening incident response, and improving release safety through progressive delivery and automated deployment analysis

Технологии и навыки

Site Reliability Engineering
Amazon Web Services
Grafana
Kubernetes
Log management tools
Prometheus
Terraform
AIOps
Incident Management (ITSM)
Load Testing
Stress testing
E
EPAM

ГрейдSenior
ЗанятостьПолная занятость
РегионАрмения
ФорматУдалённо
ИсточникСкрыто
Опубликовано26 авг.
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо

Похожие вакансии

Agentic AI DeveloperСБЕРМладший системный инженер (Ceph)Лаборатория КасперскогоРуководитель направления аналитики взысканияАк Барс БанкAI-инженерДневник.ру
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).