WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Senior Site Reliability Engineer - Fleet
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Senior Site Reliability Engineer - Fleet

Lambda·San Francisco Office (Fremont St)·4 дн назад

Senior Site Reliability Engineer - Fleet

🏢 ОфисSeniorПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Наша компания

Founded in 2012, with 500+ employees, and growing fast Our investors notably include TWG Global, US Innovative Technology Fund (USIT), Andra Capital, SGW, Andrej Karpathy, ARK Invest, Fincadia Advisors, G Squared, In-Q-Tel (IQT), KHK & Partners, NVIDIA, Pegatron, Supermicro, Wistron, Wiwynn, Gradient Ventures, Mercato Partners, SVB, 1517, and Crescent Cove We have research papers accepted at top machine learning and graphics conferences, including NeurIPS, ICCV, SIGGRAPH, and TOG Our values are publicly available:

Чем предстоит заниматься

Build and operate monitoring and alerting for cluster health — fabric, GPU, power/thermal, and job-level signals — to detect and respond to issues proactively
Remotely deploy and configure large-scale HPC clusters for AI workloads using automation wherever possible
Automate cluster lifecycle: operating systems, firmware, drivers, and networking, managed as code (Ansible, Terraform) rather than by hand
Create runbooks and automated remediations for common cluster failure modes, designed so Support and HPC Support can run them safely
Troubleshoot and resolve cluster issues across InfiniBand/RoCE, NCCL, GPU-direct, fabric, switching, and power — working closely with on-site deployment teams
Participate in on-call rotations and lead incident response for cluster-level problems
Contribute to and maintain Standard Operating Procedures, and feed clear requirements back to other engineering teams on simplification, stability, and operational efficiency

Наши требования

Experience with machine learning / deep learning frameworks (PyTorch, TensorFlow) and benchmarking tools (DeepSpeed, MLPerf)
Knowledge of containerization and orchestration technologies (e.g., Docker, Kubernetes)
Experience building and/or operating HPC resources
Depth in the NVIDIA hardware and firmware ecosystem
Experience with data center power and thermal design
Background in chaos engineering or similar reliability testing methodologies
Understanding of compliance frameworks (SOC 2, ISO 27001, etc.)
Salary Range Information
7+ years of experience in Site Reliability Engineering, HPC Engineering, DevOps, or a similar role
Have a strong understanding of modern AI infrastructure, from GPU architectures to hardware performance optimization
Strong understanding of Linux-based systems in a distributed environment
Are experienced configuring and troubleshooting InfiniBand (IB), RoCE, CLOS fabrics, 100GbE, Ethernet/switching, GPU-direct, and NCCL environments
Solid understanding of Python and Go, with experience working with SWE teams to improve internal tooling
Experience with monitoring and alerting tools (e.g., Prometheus, Grafana, Clickhouse)
Proficiency in automation and configuration management tools (e.g., Ansible, Terraform)
Have excellent problem-solving and troubleshooting skills and an innate attention to detail
Passion for continuous improvement and innovation

Мы предлагаем

Health, dental, and vision coverage for you and your dependents
Wellness and commuter stipends for select roles
401k Plan with 2% company match (USA employees)
Flexible paid time off plan that we all actually use

Дополнительно

The annual salary range for this position has been set based on market data and other factors. However, a salary higher or lower than this range may be appropriate for a candidate whose qualifications differ meaningfully from those listed in the job description
L
Lambda
San Francisco Office (Fremont St)

ГрейдSenior
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано4 дн назад
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).