WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Senior Site Reliability Engineer - Core Cloud Platform
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Senior Site Reliability Engineer - Core Cloud Platform

Lambda·San Francisco Office (Fremont St)·29 июля

Senior Site Reliability Engineer - Core Cloud Platform

🏢 ОфисSeniorПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Наша компания

Founded in 2012, with 500+ employees, and growing fast Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG Our values are publicly available:

Чем предстоит заниматься

Lambda’s Core Cloud Platform powers compute provisioning and infrastructure orchestration across our physical data centers. We are looking for a Senior Site Reliability Engineer to improve the reliability, scalability, and operational maturity of these systems as Lambda’s fleet and customer base grow
You will work across Kubernetes, infrastructure automation, observability, deployment systems, and incident response to build a resilient foundation for customer AI workloads
Operate and scale critical platform services across Lambda’s data centers
Improve the reliability of compute provisioning, Instance lifecycle, and regional orchestration systems
Build monitoring, alerting, and tracing for service health, provisioning latency, and customer-impacting failures
Define SLIs, SLOs, error budgets, and operational readiness standards
Automate detection and remediation of configuration drift, failed workflows, and orphaned resources
Build safe deployment, rollback, and disaster recovery workflows using infrastructure as code and GitOps
Design fault-isolation mechanisms that reduce blast radius and prevent cascading failures
Lead production incident response, postmortems, and durable corrective actions
Partner with Compute, Networking, Storage, Security, and Support teams
Participate in on-call and improve its sustainability through automation and better tooling
Mentor engineers and raise the reliability bar across the organization

Наши требования

Experience with AI infrastructure, GPU platforms, or high-performance computing
Experience operating distributed systems across multiple regions or data centers
Experience with Kubernetes controllers, operators, CRDs, admission control, or scheduler extensions
Experience with etcd performance, backup, restore, or disaster recovery
Experience with Linux systems, container runtimes, cgroups, storage, or networking
Experience with chaos engineering, fault injection, or automated remediation
Experience with Kubernetes RBAC, OIDC, workload identity, or certificate management
Familiarity with SOC 2, ISO 27001, or similar compliance frameworks
Salary Range Information
Have 7+ years of experience in site reliability, infrastructure, distributed systems, or production software engineering
Have deep experience operating Kubernetes in production
Understand Kubernetes architecture, scheduling, networking, resource management, upgrades, and common failure modes
Have experience with physical data centers, private cloud, hybrid cloud, or environments without full reliance on managed services
Are proficient with Terraform or similar infrastructure-as-code tools
Have built CI/CD or GitOps workflows using tools such as Argo CD, Flux, Helm, or Kustomize
Have experience with observability platforms such as OpenTelemetry, Prometheus, Grafana, or Datadog
Can build production-quality tooling in Go, Python, or a similar language
Understand distributed systems concepts including consistency, retries, idempotency, backpressure, and partial failure
Have experience defining and operating against SLIs and SLOs
Can lead effectively during high-severity incidents
Approach recurring operational issues as engineering and automation problems
Communicate clearly and work effectively across teams
Bring strong ownership, sound judgment, and low ego

Мы предлагаем

Health, dental, and vision coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible paid time off plan that we all actually use

Дополнительно

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description
L
Lambda
San Francisco Office (Fremont St)

ГрейдSenior
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано29 июля
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).