We’re seeking a Principal Network Engineer (L7) focused on owning and evolving large-scale, RoCEv2 data center networks powering next generation AI and ML infrastructure
You’ll work closely with our network architect and infrastructure leadership to define how the network is designed, implemented, and operated at scale, keeping over 8,000 GPUs burring today and scaling to cluster sizes reaching over 100,000 GPUs. You will be responsible for the architectural decisions that determine performance, reliability, and operational sanity at scale
You’ll remain hands-on with high-speed optics, switching, routing, and congestion management in production clusters, while also setting the standards, patterns, and tooling other engineers build and operate against
As a Principal Engineer, you own end-to-end network architecture, make high-impact design decisions, and set technical direction across teams, with clear examples of systems you’ve defined and scaled
Define, evolve, and standardize large-scale RoCEv2 data center networks supporting AI and ML clusters from thousands to 100,000+ GPUs
Set and validate congestion management strategy across RDMA fabrics, including PFC, ECN, and DCQCN, based on real production behavior
Establish automation, validation, and observability patterns that prevent misconfiguration and eliminate manual operational work
Act as the technical escalation point for complex failures, scaling limits, and architectural tradeoffs in always-on, multi-tenant environments