B

Production Engineer - Applied Machine Learning

Bytedance
San Jose, California, United States
🏢 Офис
Без опыта
Полная занятость
Продуктовая
Tech
Описание вакансии
Обязанности
Build reliability mechanisms for SLO SLA observability and alerting
Create automated inspections and pre flight checks
Develop CI/CD canary releases and auto rollback
Forecast capacity and manage elastic auto scaling
Govern GPU CPU storage and network resources
Handle fault diagnosis incident response and post mortems
Implement auto healing disaster recovery and incident reviews
Manage production stability for training and inference systems
Manage quotas cost attribution and performance tuning
Orchestrate AML pipelines

Технологии: Alerting, Auto Scaling, Auto rollback, Auto-healing, CI/CD, Canary Releases, Capacity forecasting, Cost attribution, Disaster Recovery, Distributed Training, GPU, Incident Response, Inference Serving, Kubernetes, Model Inference, NoSQL, Observability, Online Inference, Online Inference Serving, Orchestration, Parameter Server, Performance Tuning, Postmortems, Quota Management, Resource Governance, SLA, SLO, Scheduling

AI-помощник
ИсточникСкрыто
Опубликовано2 дн назад
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение — прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).