B
Production Engineer - Applied Machine Learning
Bytedance
San Jose, California, United States
🏢 Офис
Без опыта
Полная занятость
Продуктовая
Tech
Описание вакансии
Обязанности
•Build reliability mechanisms for SLO SLA observability and alerting
•Create automated inspections and pre flight checks
•Develop CI/CD canary releases and auto rollback
•Forecast capacity and manage elastic auto scaling
•Govern GPU CPU storage and network resources
•Handle fault diagnosis incident response and post mortems
•Implement auto healing disaster recovery and incident reviews
•Manage production stability for training and inference systems
•Manage quotas cost attribution and performance tuning
•Orchestrate AML pipelines
Технологии: Alerting, Auto Scaling, Auto rollback, Auto-healing, CI/CD, Canary Releases, Capacity forecasting, Cost attribution, Disaster Recovery, Distributed Training, GPU, Incident Response, Inference Serving, Kubernetes, Model Inference, NoSQL, Observability, Online Inference, Online Inference Serving, Orchestration, Parameter Server, Performance Tuning, Postmortems, Quota Management, Resource Governance, SLA, SLO, Scheduling