15+ years of experience in infrastructure, networking, or production engineering — with meaningful time at companies operating at internet scale (cloud providers, CDNs, large-scale social/media platforms, or similar)
Strong systems fundamentals: Linux, distributed systems, storage, compute scheduling — you understand the full stack from hardware up
Hands-on data center experience: you've done physical infra, understand power and thermal constraints, and can reason about reliability at the facility level, not just the server level
The ability to write code — not necessarily full-time, but enough to automate what shouldn't be manual, instrument what isn't observable, and build tooling your team will actually use
Excellent analytical and problem-solving skills, including the ability to synthesize ambiguous customer and system signals into clear problem statements and designs
Strong incident command: you lead calmly under pressure, communicate clearly during outages, and run blameless retrospectives that actually improve systems
Deep networking expertise: BGP, OSPF, ECMP, load balancing, and low-latency network design in production — you can debug a routing issue and design a fabric, sometimes in the same incident
Experience with HPC infrastructure: GPU cluster operations, job schedulers (Slurm, Kubernetes), high-bandwidth interconnects (InfiniBand, RoCE)
Prior principal or staff IC role where you influenced org-level technical strategy, not just project-level execution
Exposure to sustainability-focused or energy-constrained compute environments