WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Что спрашивают
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыТренды рынкаОфферыВопросы с собеседованийВозможностиИИ-инструментыПроверка резюме без входаСоветыСоздать резюмеТренировка интервьюИгра «Путь джуна»
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Site Reliability Engineer
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Site Reliability Engineer

xAI·Memphis, Tennessee; Southaven, Mississippi·3 дн назад

Site Reliability Engineer

🏢 ОфисMiddleПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Наша компания

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. Al

О роли

As a Site Reliability Engineer focused on campus reliability, you will design what the campus watches and trusts, technically command cross-discipline SEVs, and build the guardrails that make the next incident smaller. You are the connective tissue across compute, network, storage, power, and cooling. This role demands calm incident leadership, fleet-scale observability judgment, and the ability to drive reliability work across software and facility boundaries

Чем предстоит заниматься

Own monitoring architecture and signal quality: what we alert on, suppress, and trust. Consume NOC noise-disposition feedback to drive suppression and redesign. Treat alert noise as a design failure, not an operator failure
Provide SEV command support: technical incident leadership, bridge coordination with the NOC, and timeline and severity hygiene
Run blameless postmortems and drive corrective actions to closed, not filed
Lead cross-functional reliability projects spanning compute, network, storage, and facility signal boundaries
Build and maintain playbooks, run game days, and keep cross-discipline dependency maps current. Own runbook quality jointly with the NOC (SRE designs; NOC operates and corrects)
Define error budgets and availability objectives at campus and service boundaries as adopted by the business
Participate in on-call rotations and incident response for SEV-class events in the Memphis / Southaven data center campus

Наши требования

Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience)
5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments
Proven large-scale incident command experience and calm technical leadership on a bridge
Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality
Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry
Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner
Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them
Excellent problem-solving skills with a data-driven approach to reliability engineering
Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering
Experience in AI/ML infrastructure or supercomputing environments
Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries
Experience running game days, dependency mapping, and closed-loop corrective action programs
Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry
Prior work in a fast-paced startup or tech company like SpaceXAI
x
xAI
Memphis, Tennessee; Southaven, Mississippi

ГрейдMiddle
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано3 дн назад
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).