L

Staff/Principal DevOps Engineer, AI Inference

Lila Sciences
Cambridge, MA USA
🏢 Офис
Senior
Полная занятость
Продуктовая
Tech
Описание вакансии
Обязанности
Build CI CD pipelines for model artifacts and benchmarking
Build intelligent request routing and load balancing
Configure autoscaling for inference compute capacity
Create infrastructure as code for EKS and GPU clusters
Deploy inference pipelines with canary rollouts A B testing and rollback
Design GPU accelerator infrastructure on Kubernetes
Implement model serving platforms
Implement observability for latency and GPU utilization
Optimize inference performance and token throughput
Optimize request batching and caching
Условия
Commuter benefits
Company Subsidized Lunch Program
Company holidays
Dental insurance
Educational assistance
Flexible time off
Life and disability insurance
Medical insurance
Paid parental leave
Vision insurance

Технологии: AWS IAM, AWS S3, Amazon EC2, Amazon EKS, Amazon VPC, Autoscaling, CI/CD, CUDA, CUDA Driver, Continuous batching, Docker, EFA, Helm, Inference Server, Infrastructure as Code, KV cache, Kubernetes, Latency profiling, Load Balancing, Monitoring, NCCL, NVIDIA GPUs, Network interconnect, Pipeline parallelism, PrivateLink, Python, Quantization, SLI, SLO, Speculative decoding, Tensor Parallelism, Terraform, Token Throughput, Triton Inference, Triton Inference Server, VLLM, “as-code”

AI-помощник
ИсточникСкрыто
Опубликовано2 дн назад
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение — прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).