Bachelor's degree in Systems Engineering, Computer Science, Electrical Engineering, or a related field (or equivalent experience)
5+ years of experience in site reliability, systems engineering, or large-scale production operations, preferably in high-performance computing or data center environments
Proven large-scale incident command experience and calm technical leadership on a bridge
Demonstrated monitoring and observability design at fleet or campus scale, including alert hygiene, suppression, and signal quality
Experience working across at least two of: compute, network, storage, power, and cooling / facilities telemetry
Experience writing and operating playbooks or runbooks with a 24/7 operations or NOC partner
Proficiency in scripting (Python, Bash) for automation and analysis, plus general experience in at least one systems language (C, C++, Java, Go, Rust, or similar). Not required to be expert in all of them
Excellent problem-solving skills with a data-driven approach to reliability engineering
Ability to work collaboratively with cross-functional teams, including NOC, data center operations, and infrastructure engineering
Experience in AI/ML infrastructure or supercomputing environments
Hands-on definition and use of SLOs, SLIs, and error budgets at service or campus boundaries
Experience running game days, dependency mapping, and closed-loop corrective action programs
Familiarity with data center hardware and plant signals (servers, GPUs, networking, power, cooling) in addition to software telemetry
Prior work in a fast-paced startup or tech company like SpaceXAI