WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Principal Site Reliability Engineer (SRE)
ВердиктОписаниеИнструментыКомпанияПохожие
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Principal Site Reliability Engineer (SRE)

EPAM·29 авг.

Principal Site Reliability Engineer (SRE)

🌍 УдалённоHeadПолная занятостьАутсорс
Зарплата не указана
44
Есть о чём спросить
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Чем предстоит заниматься

Build self-service developer tooling, golden paths, and automated environment provisioning pipelines so development teams can deploy microservices safely and independently
Provision, harden, and manage production-grade Amazon EKS clusters using modular Terraform, Karpenter autoscaling, and GitOps delivery patterns
Design and establish the organization's incident response model, on-call escalation policies, and blameless post-mortem processes from the ground up
Lead the enterprise implementation of Datadog, including APM, distributed tracing, and custom metrics, while defining meaningful Service Level Objectives (SLOs) and actionable alerting rules
Architect GitHub Actions CI/CD workflows supporting zero-downtime progressive delivery strategies such as Canary and Blue/Green releases, along with automated health verification
Partner with the Solution Architect to implement robust cloud network segregation, including VPCs, transit gateways, IRSA, and ingress security boundaries

Наши требования

7+ years of experience as a Site Reliability Engineer, Platform Engineer, or similar role focused on cloud-native infrastructure
Expertise in AWS, Amazon EKS, and Kubernetes in production environments
Proficiency in Terraform, Karpenter, and GitOps delivery patterns
Background in designing incident response frameworks, on-call models, and blameless post-mortem processes
Skills in Datadog for APM, distributed tracing, and custom metrics, along with defining SLOs and alerting rules
Competency in GitHub Actions for CI/CD workflows and progressive delivery strategies such as Canary and Blue/Green releases
Understanding of cloud network architecture, including VPCs, transit gateways, IRSA, and ingress security boundaries
Familiarity with Kotlin backend and React Native mobile development ecosystems to effectively support developer enablement
English proficiency at B2 level or higher

Дополнительно

We are seeking an experienced Principal Site Reliability Engineer (SRE)
To architect, build, and operate the foundational platform infrastructure for a greenfield, cloud-native platform on AWS. This role is 100% focused on proactive platform engineering, developer enablement, and reliability architecture, not daily ticket handling or manual operations. Operating as an individual contributor, you will design and implement an Internal Developer Platform (IDP) to empower Kotlin backend and React Native mobile development teams, while establishing the organization's incident response frameworks, on-call models, and observability standards from scratch

Технологии и навыки

Platform Engineering
Amazon Web Services
Deployment Strategies
Infrastructure as Code development and maintenance
Network Architecture
Site Reliability Engineering
Argo CD
Java Microservice Infrastructure
Observability and troubleshooting in distributed systems
TypeScript
E
EPAM

ГрейдHead
ЗанятостьПолная занятость
РегионКазахстан
ФорматУдалённо
ИсточникСкрыто
Опубликовано29 авг.
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо

Похожие вакансии

SAP Basis ArchitectProДоступна только зарегистрированнымSAP BTP ArchitectProДоступна только зарегистрированнымSenior DevOps/Platform EngineersEPAMLead AI Python EngineerEPAM
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).