WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Что спрашивают
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыТренды рынкаОфферыВопросы с собеседованийВозможностиИИ-инструментыПроверка резюме без входаСоветыСоздать резюмеТренировка интервьюИгра «Путь джуна»
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Senior Software Engineer - AI Infrastructure Performance Insights & Observability
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Senior Software Engineer - AI Infrastructure Performance Insights & Observability

CoreWeave·Sunnyvale, CA / Bellevue, WA·13 авг.

Senior Software Engineer - AI Infrastructure Performance Insights & Observability

≈ от 200 000 ₽наша оценка по вакансиям этой роли и грейда, у работодателя вилка не указана
🏢 ОфисSeniorПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан
Написать напрямуюВ тексте вакансии есть контакт: письмо уйдёт человеку, а не в систему подбора

Наша компания

CoreWeave is The Essential Cloud for AI™. Built for pioneers by pioneers, CoreWeave delivers a platform of technology, tools, and teams that enables innovators to build and scale AI with confidence. Trusted by leading AI labs, startups, and global enterprises, CoreWeave combines superior infrastructure performance with deep technical expertise to accelerate breakthroughs and turn compute into capability. Founded in 2017, CoreWeave became a publicly traded company (Nasdaq: CRWV) in March 2025. About this role

О роли

We're looking for a Senior Engineer to be a driving force on CoreWeave's Benchmarking & Performance team, with a focus on building the performance insights and observability systems that make our AI infrastructure legible at every layer, from individual GPUs and NVLink/InfiniBand fabrics up through distributed training and inference workloads. You will own how we detect, diagnose, and surface perf

Чем предстоит заниматься

Performance Insights & Observability - Design and build the systems that continuously assess AI infrastructure health and performance: GPU utilization and efficiency, interconnect (NVLink, InfiniBand, RoCE) fabric behavior, distributed training and inference throughput, and hardware degradation signals. Build the detection and diagnosis logic that surfaces anomalies and regressions before they become incidents, not just dashboards that report on them after the fact
Time-Series & Metrics Infrastructure - Own and extend our time-series database (TSDB) layer as the backbone of real-time observability. Write and optimize PromQL/MetricsQL queries that power alerting, anomaly detection, and trend analysis across thousands of GPUs and hundreds of benchmark runs. Bridge streaming metrics and batch-analytical workloads so engineers get sub-second answers during live incidents and analysts get complete historical context for root cause work
Fabric & GPU Telemetry - Build and validate the pipelines and metrics that make network fabric and GPU-level behavior observable and comparable across racks, clusters, and hardware generations, including gray failure detection, congestion and error-rate signals, and health scoring that holds up under audit
Data Lake Architecture (in support of insight work) - Design and build the performance data lake that underpins the above: table formats (Apache Iceberg, Parquet, Avro), hot/cold tiering, and schema evolution for latency distributions, throughput metrics, GPU utilization, cost-per-token, and hardware health signals. This is foundational infrastructure, not the end product
Query Optimization & Performance - Profile and tune query engines against columnar and time-series stores so that the observability layer meets its own strict P99 latency and freshness SLAs. Benchmark the benchmarking infrastructure itself
BI & Reporting (secondary) - Where needed, build self-service views (Grafana, Looker, or similar) for engineers, product managers, and executives, but as a downstream output of the insight and observability work above, not the primary deliverable

Наши требования

5+ years of experience building distributed systems, observability platforms, or performance engineering tooling, ideally for infrastructure or ML systems rather than general-purpose BI
Strong coding in Python or Go (C++ a plus) and deep familiarity with networked systems, GPU infrastructure, and performance analysis
Hands-on experience with Kubernetes at production scale, CI/CD, and observability stacks (Prometheus, Grafana, OpenTelemetry) used to monitor and diagnose infrastructure, not just report on it
Working knowledge of time-series databases and fluency in PromQL or MetricsQL for building real-time alerting and anomaly detection, not only historical dashboards
Familiarity with data lake architectures and modern table formats (Iceberg, Parquet, Avro) sufficient to support an insights platform, though this is not the primary skill this role is hiring for
Comfortable working close to hardware and workload behavior: GPU utilization patterns, interconnect fabric health, distributed training/inference performance characteristics
Strong communicator comfortable collaborating with cross-functional teams and external partners
Experience with time-series databases, LSM-based storage engines, or custom telemetry pipelines
Experience running MLPerf submissions or similar large-scale audited benchmarks
Contributions to OSS projects such as Apache Iceberg, Apache Spark, Trino, llm-d, vLLM, or PyTorch
Direct experience benchmarking or monitoring large GPU fleets or multi-region clusters
Experience with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies
Familiarity with data cataloging, lineage tools, or data governance frameworks
Why CoreWeave?
Own the data that tells the truth about performance. You'll build the analytical foundations the industry relies on to measure—and achieve—state-of-the-art results, working hand-in-hand with world-class partners and communities. If you love turning vast streams of telemetry into clear, trusted insights that move markets and drive engineering excellence, we'd love to talk
At CoreWeave, we work hard, have fun, and move fast! We're in an exciting stage of hyper-growth that you will not want to miss out on. We're not afraid of a little chaos, and we're constantly learning. Our team cares deeply about how we build our product and how we work together, which is represented through our core values
Be Curious at Your Core
Act Like an Owner
Empower Employees
Deliver Best-in-Class Client Experiences
Achieve More Together
We support and encourage an entrepreneurial outlook and independent thinking. We foster an environment that encourages collaboration and provides the opportunity to develop innovative solutions to complex problems. As we get set for take off, the growth opportunities within the organization are constantly expanding. You will be surrounded by some of the best talent in the industry, who will want to learn from you, too. Come join us!

Мы предлагаем

In addition to a competitive salary, we offer a variety of benefits to support your needs. The benefits below reflect our US-based offerings for full-time employees; for roles in other locations, benefits vary and are shared during the hiring process. These include
Medical, dental, and vision insurance - 100% paid for by CoreWeave
Company-paid Life Insurance
Voluntary supplemental life insurance
Short and long-term disability insurance
Flexible Spending Account
Health Savings Account
Tuition Reimbursement
Ability to Participate in Employee Stock Purchase Program (ESPP)
Mental Wellness Benefits through Spring Health
Family-Forming support provided by Carrot
Paid Parental Leave
Flexible, full-service childcare support with Kinside
401(k) with a generous employer match
Flexible PTO
Catered lunch each day in our office and data center locations
A casual work environment
A work culture focused on innovative disruption
California Applicants
California Consumer Privacy Act

Дополнительно

The base salary range for this role is $182,000 to $242,000. The starting salary will be determined based on job-related knowledge, skills, experience, and market location. We strive for both market alignment and internal equity when determining compensation. In addition to base salary, our total rewards package includes a discretionary bonus, equity awards, and a comprehensive benefits program (all based on eligibility)
The range we’ve posted represents the typical compensation range for this role. To determine actual compensation, we review the market rate for each candidate which can include a variety of factors. These include qualifications, experience, interview performance, and location
C
CoreWeave
Sunnyvale, CA / Bellevue, WA

ГрейдSenior
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано13 авг.
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).