Наши требования
5–9 years in Infrastructure Engineering, Platform Engineering, or SRE, with a specific focus on high-performance computing or large-scale AI stacks. Proven track record of managing complex production environments where system reliability is mission-critical
Kubernetes Expert: Deep experience in cluster administration and scheduler internals; comfortable reading/modifying controller code
AI/GPU Infrastructure Specialist: Proficient in orchestrating GPU workloads and diagnosing training job failures using ROCm or CUDA
Network Pathologist: Skilled in RDMA/RoCEv2, SRIOV, and BGP; capable of interpreting switch telemetry to identify silent packet drops
Linux Power User: Expert in kernel networking, hugepages, and cgroups; able to debug at the OS layer when applications are silent
Builder Mindset: Proficient in Python and Ansible; capable of writing custom diagnostic tools to automate remediation
Executive Communicator: Strong technical rigor when presenting findings to VPs of Engineering, maintaining trust while delivering difficult updates
Prior experience in a customer-facing engineering role (e.g., Solutions Engineering, Technical Support Engineering)
Experience in high-uptime environments where 24/7/365 availability is required