Наши требования
3–5 years of experience building distributed systems, high-performance computing components, or cloud services
Strong programming skills in Python or Go (C++ a plus) with understanding of networked systems and performance fundamentals
Hands-on experience with Kubernetes in production environments plus familiarity with CI/CD and observability tools (e.g., Prometheus, Grafana, OpenTelemetry)
Exposure to performance-critical GPU systems (CUDA, NCCL, NVLink/PCIe, memory bandwidth) or model-serving stacks (llm-d, vLLM, TensorRT-LLM, Megatron-LM)
Effective communicator comfortable working cross-functionally
Experience with time-series databases, LSM-based storage engines, or custom data pipelines
Familiarity with MLPerf or other large-scale benchmarking frameworks
Contributions to OSS projects such as llm-d, vLLM or PyTorch
Exposure to benchmarking GPU clusters or multi-region environments
Background working with CUDA kernels, NCCL/SHARP, RDMA/NUMA, or GPU interconnect topologies