WorkaemКарьерная платформа
  • Вакансии
  • Компании
  • Зарплаты
  • Офферы
  • Сервисы
  • Блог
  • Работодателям
Workaem

Карьерная платформа для IT-специалистов: вакансии напрямую с карьерных страниц 300+ компаний, из телеграм-каналов, с международных площадок и от работодателей напрямую. Разбор условий, детектор мёртвых вакансий, AI-инструменты для резюме. Базовые функции бесплатны.

Подпишись, присылаем лучшие вакансии недели
Или читай канал в телеграме
Соискателям
Все вакансииЗа границейУдалёнка в долларахКомпании с РУ основателямиЗарплатыОфферыВозможностиСоветыСоздать резюмеТренировка интервью
По технологиям
Вакансии PythonВакансии JavaScriptВакансии ReactВакансии JavaВакансии GoВакансии Docker
По профессиям
РазработкаДизайнQA / ТестированиеАналитикаProduct / Project ManagerМаркетинг
Работодателям
Разместить вакансиюТарифыБаза кандидатовСвязаться с нами
Кабинет
РегистрацияВойтиЛичный кабинетМои откликиСохранённыеУведомления
Компания
О проектеПредложенияКонтактыБлогКонфиденциальностьУсловия использования
© 2026 Workaem. Все права защищены.КонфиденциальностьУсловияОферта
Made by IT, for IT 💛
Network Engineer (Supercomputer Infrastructure)
ВердиктОписаниеИнструментыКомпания
  1. Главная
  2. /
  3. Вакансии
  4. /
  5. Network Engineer (Supercomputer Infrastructure)

xAI·Memphis, Tennessee; Southaven, Mississippi·3 дн назад

Network Engineer (Supercomputer Infrastructure)

🏢 ОфисMiddleПолная занятость
Зарплата не указана
Вилки нет, про деньги придётся договариваться с нуля.
Нажмите на сигнал, чтобы увидеть, на чём он основан

Наша компания

SpaceXAI’s mission is to create AI systems that can accurately understand the universe and aid humanity in its pursuit of knowledge. Our team is small, highly motivated, and focused on engineering excellence. This organization is for individuals who appreciate challenging themselves and thrive on curiosity. We operate with a flat organizational structure. All employees are expected to be hands-on and to contribute directly to the company’s mission. Leadership is given to those who show initiative and consistently deliver excellence. Work ethic and strong prioritization skills are important. Al

О роли

SpaceXAI is looking for an exceptional network engineer with experience in mission-critical, large-scale production environments to support the design, build-out, and operation of networks that power our AI supercomputer campuses. As a member of the Supercomputer Infrastructure / Network Engineering team, you will provide design and operational support for the fabrics used by GPU training and inference clusters, site operations, automation and controls, and facilities teams. The ideal candidate thrives in intense, high-flux environments, brings a strong sense of urgency balanced with operational excellence, communicates clearly, and demonstrates high technical acumen

Чем предстоит заниматься

Design and implement highly available, low-latency, high-bandwidth networks, carefully balancing routing, congestion control, and redundancy technologies for AI training fabrics, inference front-ends, storage, and site/OT networks
Design and maintain supercomputer data center and campus networks in accordance with company network standards. Collaborate with adjacent infrastructure, compute, storage, SiteOps, and enterprise teams
Evaluate, procure, and deploy network hardware including data-center class switches, NICs, firewalls, optical multiplexers, and related appliances supporting 400G/800G and beyond
Contribute to maturing network automation tooling; implement configuration analysis, linting, validation, and scalable deployment frameworks (GitOps / IaC)
Plan and coordinate network change windows with stakeholders to perform software updates, hardware refreshes, cluster expansions, and general maintenance (including evenings and weekends when required by compute schedules)
Troubleshoot and resolve network-related issues affecting cluster health and job performance; publish root cause analysis (RCA) documentation and host retrospectives
Provide direct networking support during cluster bring-up, expansion, and production training/inference campaigns; serve as on-call or networking responsible engineer during operations
Proactively tailor network monitoring and telemetry (fabric health, congestion, packet loss, NCCL/collective performance) so issues are detected before they impact training or inference
Continuously create and update network documentation, including architecture overviews, design drawings, fiber/cable plant records, and operational procedures
Collaborate with cross-functional teams to identify and resolve potential design issues, especially systemic or cascading failure modes and false redundancy in AI fabrics and site networks
Perform job walks with customers, vendors, and contractors to gather requirements and produce implementation plans for new halls, rows, and campus interconnects
Ensure networks are configured and maintained in compliance with industry and cybersecurity standards (e.g., ITAR, ISO, NIST), with particular attention to segmentation between compute fabrics, storage, OT/controls, and corporate networks

Наши требования

Bachelor’s degree in computer science, computer engineering, or other STEM discipline and 3+ years of professional network engineering experience
OR 5+ years of professional network engineering experience in lieu of a degree
Extensive hands-on experience designing, deploying, supporting, and troubleshooting Layer 2 and Layer 3 networks in latency-sensitive and/or industrial / data-center environments
Functional experience with multiple network vendors in production or lab environments
Experience with GitOps and Infrastructure as Code frameworks, both as a user and contributor
Strong understanding of the OSI model and network standards
Hands-on experience with Cisco, Arista, Juniper, and/or NVIDIA Spectrum-X data-center class switches
Experience with RoCEv2 Ethernet AI/HPC fabrics; InfiniBand experience is a plus
Working knowledge of AI training and inference traffic patterns and how they behave on the network (collectives, congestion, ECMP, adaptive routing). Familiarity with NCCL is a plus
Experience with WDM and large-scale single-mode / multimode fiber plants, including OTDR and acceptance testing
Experience with switch port security, network segmentation, QoS, multicast, and redundancy protocols
Familiarity with network monitoring and Layer 1 test tools; experience building operational telemetry and dashboards
Proficiency in scripting (Bash / PowerShell / Python) and automation frameworks (Terraform, Ansible, etc.)
Linux and Windows system administration experience professionally or from labs
Industry-standard certifications such as CCNA or CCNP
Experience supporting real-time systems, industrial control / OT networks, or high-reliability environments in data center, energy, aerospace, defense, or similar industries
Excellent communication skills with internal and external customers, vendors, and management in both formal and informal settings
Ability to pass applicable background checks for site access
Ability to work in tight quarters; physical dexterity is necessary to perform job functions
Availability for extended hours and/or weekends as the schedule varies with cluster build-out and operational needs; flexibility is required
Ability to provide 24x7 on-call support in emergency situations and participate in an after-hours on-call rotation
Willingness to travel (up to 20%) between supercomputer campuses and related sites
Ability to lift 30 lbs
Ability to work at heights
Ability to drive (active valid driver’s license)
x
xAI
Memphis, Tennessee; Southaven, Mississippi

ГрейдMiddle
ЗанятостьПолная занятость
РегионНе Россия
ФорматОфис
ИсточникСкрыто
Опубликовано3 дн назад
Все вакансии компании

AI-помощник

под эту вакансию
Войди, чтобы AI оценил твоё соответствие вакансии и написал сопроводительное письмо
Мы против мошенников на площадке: если тебя просят заплатить, продиктовать код или установить непонятное приложение, прекращай общение и сразу пиши нам (чат с основателем или форма обратной связи).