Наши требования
10+ years building infrastructure-layer systems at scale — fleet management, distributed control planes, scheduler internals, or hardware lifecycle automation. This is a systems-builder role, not a consumer of managed cloud services
Deep experience with distributed systems design: consensus, state reconciliation, closed-loop automation, and systems that make autonomous decisions against live production infrastructure
Hands-on fluency with GPU/HPC infrastructure — GPU health telemetry, NVLink/InfiniBand/RoCE fabrics, thermal and power behavior at the hardware level
Track record of designing and shipping large-scale observability or telemetry platforms that correlate signals across compute, network, and storage layers
Comfort operating in ambiguity and defining the architecture and standards for a system that doesn't exist yet — this is a 0→1 charter, not a maintenance role
Strong software engineering fundamentals in at least one systems language (Go, Rust, C++, or similar) and the judgment to know when to build vs. adopt existing tooling
Experience applying ML/statistical methods to noisy operational telemetry (failure prediction, anomaly detection) is a strong plus
Prior exposure to zero-trust or policy-based multi-tenancy architectures is a plus but not required