Чем предстоит заниматься
Keep real-time inference services healthy through alerting, live debugging and fixing of production incidents, rollbacks and postmortems, as part of a shared 24x7 on-call rotation
Build and improve AWS infrastructure, including infrastructure as code, CI/CD pipelines, Kubernetes, Docker, autoscaling, monitoring, load testing and cost control
Care for the GPU fleet through NVIDIA driver and CUDA upgrades, node provisioning and debugging, and capacity management
Automate user access, including SSH credentials, service accounts and developer environments for the science team
Run real-time production data pipelines, including streaming ingestion into an online feature store and serving features to low-latency endpoints
Orchestrate batch jobs with Airflow or Databricks Workflows
Partner with applied scientists to productionize their prototypes and keep them healthy
Write and review production Python, including validating AI-generated code