Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience
OR 5+ years of professional network engineering experience in lieu of a degree
Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / data-center environments
Functional experience with multiple network vendors in production or lab environments
Experience with GitOps and Infrastructure as Code frameworks, both as a user and contributor
Strong understanding of the OSI model and network standards
Hands-on experience with Cisco, Arista, Juniper, and/or NVIDIA Spectrum-X data-center class switches
Experience with RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience is a plus
Working knowledge of AI training and inference traffic patterns and how they behave on the network (collectives, congestion, ECMP, adaptive routing). Familiarity with NCCL is a plus
Experience with WDM and large-scale single-mode / multimode fiber plants, including OTDR and acceptance testing
Experience with switch port security, network segmentation, QoS, multicast, and redundancy protocols
Familiarity with network monitoring and Layer 1 test tools; experience building operational telemetry and dashboards
Proficiency in scripting (Bash / PowerShell / Python) and automation frameworks (Terraform, Ansible, etc.)
Linux and Windows system administration experience professionally or from labs
Industry-standard certifications such as CCNA or CCNP
Experience supporting real-time systems, industrial control / OT networks, or high-reliability environments in data center, energy, aerospace, defense, or similar industries
Excellent communication skills with internal and external customers, vendors, and management in both formal and informal settings
Ability to pass applicable background checks for site access
Ability to work in tight quarters; physical dexterity is necessary to perform job functions
Availability for extended hours and/or weekends as the schedule varies with cluster build-out and operational needs; flexibility is required
Ability to provide 24x7 on-call support in emergency situations and participate in an after-hours on-call rotation
Willingness to travel (up to 20%) between supercomputer campuses and related sites
Ability to lift 30 lbs
Ability to work at heights
Ability to drive (active valid driver’s license)