Дополнительно
Serve as highest-level escalation point for complex P1/P0 incidents
Lead cross-functional root cause investigations involving compute, networking (IB/RDMA/RoCE), storage, and orchestration layers
Partner with SRE, Software teams (Storage, Networking, Compute, K8) to design systemic fixes rather than recurring workarounds
Design and improve node validation, burn-in processes, performance baselining, and release readiness
Influence Kubernetes architecture, workload orchestration (Slurm, Terraform), and AI/ML cluster stability
Reduce MTTR and incident recurrence through structural improvements
Troubleshoot NCCL, IB, GPU driver/firmware issues, distributed training failures
Support complex AI workloads (training + inference) with performance tuning and observability improvements
Mentor P3/P4 engineers
Define SOPs and technical standards for support excellence
Partner with Enablement to raise the technical bar across the organization