Наши требования
Years of Infrastructure Experience: Minimum of 8+ years of experience working within infrastructure, SRE, or production engineering environments
Engineering Leadership Track Record: Minimum of 2+ years of experience directly leading first-line engineering teams within a high-growth neocloud, hyperscaler, or large-scale distributed environment
Non-Negotiable Coding Proficiency: Strong, hands-on software engineering fundamentals in Go, Python, C++, or a comparable systems language to build automation rather than scale through headcount
Distributed Systems Depth: Expert-level command of Linux internals, container orchestration at scale, and root-cause analysis across complex physical-to-virtual boundaries
Operational Execution Expertise: Proven track record of running tiered on-call models, establishing clear SLIs/SLOs and error budgets, and measurably reducing paging fatigue
AI Infrastructure Experience: Prior experience working at a neocloud or AI-infrastructure company operating massive GPU clusters
High-Performance Fabric Exposure: Hands-on exposure to high-performance networks (such as InfiniBand or RoCEv2) or hardware internals (including BMC, firmware qualification, and attestation)
Accelerator Domain Knowledge: Deep familiarity with the failure modes of modern AI accelerators (NVIDIA or AMD platforms) and how they manifest from DCGM counters up to a customer's training run
Scale Engineering: Experience managing step-change scaling milestones and building systems designed to absorb 10x fleet expansions