Чем предстоит заниматься
Design, build, and maintain scalable infrastructure to support the AI Workbench Platform and its underlying domain foundation models
Implement and manage Kubernetes clusters with multi-GPU scheduling capabilities to support large-scale model training and inference workloads
Develop and maintain Infrastructure as Code using Terraform to provision cloud resources reliably and reproducibly
Package, deploy, and manage applications using Helm and Kustomize across multiple environments
Collaborate with data scientists, ML engineers, and business units to enable seamless model consumption and fine-tuning workflows at scale
Ensure platform reliability, scalability, and security for both internal teams and external customers
Optimize resource utilization and cost efficiency across GPU-intensive workloads
Establish CI/CD pipelines and automation to accelerate platform delivery and model deployment
Monitor system performance and troubleshoot production issues to maintain high availability
Contribute to platform architecture decisions and best practices for MLOps at enterprise scale