WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Member of Technical Staff (AI Inference Engineer)
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Member of Technical Staff (AI Inference Engineer)

Perplexity AI·London·13 апр.

Member of Technical Staff (AI Inference Engineer)

🏢 ОфисSeniorПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Наша компания

We are looking for an AI Inference Engineer to join our growing team. We build and run the inference engine behind every Perplexity query and deploy dozens of model architectures at scale with tight latency and cost budgets. Our stack is Rust, Python, CUDA, and CuTe DSL.

Чем предстоит заниматься

New models support. Support transformer-based retrieval, text-generation, and multimodal models in our inference infrastructure, from weight loading, request scheduling and KV-cache management to support in API Gateway
GPU kernels migration to CuTe DSL. Port our in-house CUDA kernels to NVIDIA's CuTe DSL so they run on GB200 today and are portable to Vera Rubin racks tomorrow
Rust-native serving runtime. Develop our internal Rust-based inference server to solve all Python pains and keep up with rapidly growing traffic
Performance optimisation. Profile and fix bottlenecks from network ingress through continuous batching and GPU kernels interleaving
Reliability and observability. Build dashboards, alerts, and automated remediation so we catch regressions before users do. Respond to and learn from production incidents

Наши требования

3+ years of professional software engineering experience with meaningful work on ML inference or high-performance systems
Familiarity with at least one deep learning framework (PyTorch, JAX, TensorFlow)
Understanding of GPU architectures (memory hierarchy, warp scheduling, tensor cores)
Understanding of common LLM architectures and inference optimization techniques (e.g. quantization, speculative decoding, prefill-decode disaggregation)
Final offer amounts are determined by multiple factors including experience and expertise
Equity: In addition to the base salary, equity may be part of the total compensation package
ML compilers and framework internals: PyTorch internals, torch.compile, custom operators
Distributed GPU communication: NCCL, NVLink, InfiniBand, RDMA libraries, model/tensor parallelism
Low-precision inference: INT8/FP8/FP4 quantization, mixed-precision serving
Profiling and debugging tools: Nsight Compute/Systems, CUDA-GDB, PTX/SASS analysis
Container orchestration: Kubernetes, GPU scheduling, autoscaling inference workloads

Дополнительно

Deep experience with GPU programming and performance work (CUDA, Triton, CUTLASS, or similar). Any other deep systems programming experience is a plus
You understand modern LLM architectures and are able to bring them up reliably in a production environment
You've built and operated production distributed systems under real load - ideally performance-critical ones
Comfortable working across languages and layers: Rust for the serving runtime, Python for model code, CUDA/CuteDSL for kernels
You own problems end-to-end. You can read a research paper on Monday, write a kernel on Wednesday, and debug a production incident on Friday
Self-directed. You do well in fast-moving environments where the path forward isn't laid out for you
P
Perplexity AI
London

ГрейдSenior
ЗанятостьПолная занятость
РегионВеликобритания
ФорматОфис
ИсточникСкрыто
Опубликовано13 апр.
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).