WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Software Engineer, LLM Inference
ВердиктОписаниеИнструментыКомпанияПохожие
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Software Engineer, LLM Inference

Soniox·Любляна (Словения)·25 авг.

Software Engineer, LLM Inference

🏢 ОфисSeniorПолная занятостьАнглийский B2Релокация
Зарплата не указана
53
Есть о чём спросить
За такой навык на рынке платят больше. Но вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Описание вакансии

We’ll help you with the relocation process and paperwork from any country to Slovenia.

About the role

Soniox is pushing the boundaries of real-time AI, and we’re looking for an engineer to help us run large language models with exceptional speed, efficiency, and reliability at production scale.

In this role, you’ll work deep in the LLM inference stack, from vLLM scheduling and KV-cache management to CUDA kernels, distributed execution, and GPU profiling, optimizing every part of the path from request to generated token.

In this role, you will

Build and optimize our vLLM-based inference stack for low latency, high throughput, and maximum GPU utilization.
Optimize continuous batching, scheduling, prefill/decode, prefix caching, and KV-cache allocation and reuse.
Optimize or implement CUDA and Triton kernels for attention, GEMMs, sampling, normalization, and other critical model operations.
Evaluate and integrate technologies such as FlashAttention, FlashInfer, CUDA Graphs, torch.compile, speculative decoding, and quantization.
Optimize distributed inference using tensor, data, and expert parallelism, NCCL, NVLink/NVSwitch, and InfiniBand.
Work closely with researchers to bring new dense and MoE model architectures into production quickly and efficiently.

You might thrive in this role if you

Have deep hands-on experience with LLM inference, ideally working inside vLLM, SGLang, TensorRT-LLM, or similar systems, not just deploying them.
Understand TTFT, inter-token latency, throughput, continuous batching, PagedAttention, KV caching, and prefill vs. decode performance.
Are comfortable profiling GPUs and reasoning about compute, memory bandwidth, kernel launches, synchronization, and communication bottlenecks.
Have experience with CUDA, Triton, PyTorch, NCCL, and modern NVIDIA GPU architectures.
Understand Transformer internals including MHA/GQA, RoPE, KV cache, quantization, and MoE.
Have experience optimizing distributed, performance-critical systems in production.
Care deeply about performance, simplicity, and reliability, and take ownership from profiling through production deployment.

Why Soniox

You’ll help build one of the most technically advanced voice AI platforms in the world, and push LLM inference performance at every layer of the stack.

You’ll work directly with a world-class team of engineers and researchers on hard, measurable problems spanning models, GPU kernels, distributed systems, and production infrastructure.

You'll have a voice in how our technology evolves, how our company grows, and how AI transforms human communication.

Технологии и навыки

CUDA
PyTorch
Triton
Soniox
Soniox
Любляна (Словения)

ГрейдSenior
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано25 авг.
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо

Похожие вакансии

Старший DS-инженер в команду рекомендацийAvito363 000 – 524 000 ₽Data Scientist в команду АнтифродаAvito269 000 – 413 000 ₽Junior Research Scientist (Risk Team)ProДоступна только зарегистрированным2 300 – 3 500 €Разработчик AIProДоступна только зарегистрированным
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).