WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Member of Technical Staff - Reliability Engineering
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Member of Technical Staff - Reliability Engineering

Fireworks AI·San Mateo·17 авг.

Member of Technical Staff - Reliability Engineering

🏢 ОфисSeniorПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Наша компания

Fireworks is the platform for specialized intelligence, enabling companies to build, train, and serve AI models tailored to their own data, workflows, and products. Founded by the team behind PyTorch and backed by AMD, Atreides, Benchmark Capital, Index Ventures, Lightspeed, NVIDIA, Sequoia Capital, and TCV, Fireworks powers production AI with hundreds of state-of-the-art open models across text, image, embedding, audio, and multimodal workloads. Today, Fireworks is a Series D company valued at $17.5 billion, bringing together an ambitious, collaborative team that's building the future of enterprise AI.

О роли

Fireworks AI is one of the industry leaders in inference and training for open models. Open models are how the rest of the world gets to build on frontier AI without handing the keys to a single vendor, and our job is to make them fast, cheap, and dependable enough that this is a real choice. That work is systems work: GPU scheduling, kernel and runtime performance, networking, storage, Linux. We serve over 40 trillion tokens a day doing it

Чем предстоит заниматься

Reliability Engineering makes sure that platform runs dependably as it grows. You will work across cloud infrastructure, AI systems, and product teams to make sure the pieces fit together, fail gracefully, and hold up under load
Define reliability standards: SLOs, error budgets, production readiness criteria. Not written in a vacuum: you instrument the systems and read the real telemetry the numbers come from
Own the reliability toolchain: Logging and telemetry pipelines, alerting standards, failure injection, load testing, self-healing automation, and AI-assisted investigation tooling
Keep customer experience from falling through the cracks: Per-service reliability is necessary but not sufficient. A customer can hit a bad experience while every system sits inside its SLO. You make sure those failures get an owner and a fix
Own the seams: The hardest failures live between systems: retries that amplify load, timeouts that do not compose, dependencies nobody mapped. You find them before customers do and drive fixes through the teams that own them
Run incident management: Coordinate live production issues, run blameless postmortems, and track follow-ups to completion
Reduce toil: Automate repetitive operational work so growth does not turn into an unsustainable on-call load
Partner across the org: Cloud infrastructure on capacity and multi-region risk, inference and training on failure modes in the serving and training stacks, performance on zero-downtime rollouts, product and control plane on customer-facing reliability

Наши требования

Systems fundamentals: 5+ years with Linux internals, system performance troubleshooting, and networking fundamentals (TCP/IP, HTTP, gRPC)
Software engineering: 5+ years in Python, Go, C++, or Rust, writing production-grade tools and systems code
Cloud-native operations: Operating and debugging Kubernetes, Terraform, and Docker in high-throughput production
Distributed systems: High-throughput control planes, microservices, or multi-region setups
Reliability fundamentals: Fault-tolerant design, SLO/SLA management, automated failover, high-availability architecture
Influence without authority: You can get other teams to adopt a standard through credibility and useful tooling rather than mandate
Breadth over comfort: Willingness to dig into unfamiliar parts of the stack when a problem crosses boundaries
Education: Bachelor's or Master's in Computer Science, Computer Engineering, or equivalent practical experience
Observability tooling: Prometheus, Grafana, OpenTelemetry, and alerting people actually act on
GPU and ML infrastructure exposure: GPUs, inference serving, or distributed training
AI-assisted operations: Building agents or LLM-based tooling for investigation, triage, or automation
Open source background: Contributions to infrastructure, systems, or ML serving projects
Startup agility: Comfortable where pragmatism and teamwork matter more than process

Дополнительно

You own the bar. You define what "reliable" means at Fireworks: SLOs, error budgets, production readiness, on-call expectations. Then you drive adoption across engineering
You own the process and the tooling. Incident management, postmortems, observability standards, failure testing, guardrails, and automation are yours end to end
Every team owns the reliability of what they build. You make that ownership practical. Structured logging and aggregation, metrics and tracing that work the same way everywhere, alerting that routes to the right owner, dashboards that answer "why is this slow."
You choose where the leverage is. You have a wide view of the platform and the latitude to spend your time where it changes outcomes most
Solve Hard Problems: Tackle challenges at the forefront of AI infrastructure, from low-latency inference to scalable model serving
Build What’s Next: Work with bleeding-edge technology that impacts how businesses and developers harness AI globally
Ownership & Impact: Join a fast-growing, passionate team where your work directly shapes the future of AI—no bureaucracy, just results
Learn from the Best: Collaborate with world-class engineers and AI researchers who thrive on curiosity and innovation
Fireworks AI is an equal-opportunity employer. We celebrate diversity and are committed to creating an inclusive environment for all innovators
F
Fireworks AI
San Mateo

ГрейдSenior
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано17 авг.
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).