Explain complex concepts in clear, structured language appropriate to the reader’s experience level
Define unfamiliar terminology and provide context, glossaries, diagrams, screenshots, examples, or links where they improve understanding
Use layered documentation so that L1 technicians receive the instructions and safeguards they need, while experienced engineers can access deeper diagnostic detail
Include clear prerequisites, warnings, decision points, expected outputs, success criteria, rollback instructions, evidence-collection requirements, and escalation conditions
Show where commands must be executed, what successful output looks like, and what to do when the result differs from expectations
Hands-on experience in data center, server, infrastructure, production operations, or site reliability engineering
Working knowledge of Linux, server hardware, firmware, and out-of-band management technologies such as IPMI, BMC, OpenBMC, or Redfish
A demonstrated record of creating runbooks, SOPs, or troubleshooting guides that operations teams used successfully
Experience translating complex engineering investigations into safe, repeatable operational procedures
Strong organizational skills and the ability to structure incomplete or complex information
The ability to write clear, precise, and structured English for readers with different levels of technical experience
Experience validating documentation with its intended users and improving it through feedback
An understanding of document governance, including ownership, approvals, versioning, review cycles, deprecation, and archiving
The judgment to distinguish verified information from assumptions and to challenge incomplete or unclear technical input
Strong collaboration skills and the confidence to work closely with engineers, technicians, and knowledge-management stakeholders
Fluent written and spoken English
Willingness to travel regularly to data center locations
Experience with NVIDIA GPU server platforms and tools such as nvidia-smi, DCGM, dcgmi, and log-correlation tooling
Experience with HGX or other large-scale AI infrastructure platforms
Exposure to OCP-based platforms or ODM manufacturing ecosystems
Experience using Bash or Python for log collection, diagnostics, or operational automation
Experience with documentation-as-code, Git-based workflows, wikis, or large-scale knowledge-base platforms
Experience defining documentation metrics or using incident and escalation data to prioritize improvements
Experience supporting hardware-platform introductions across multiple data centers
What success looks like
L1 and L2 technicians resolve more issues safely without unnecessary escalation
Recurring incidents produce measurable improvements to the knowledge base
Documentation remains accurate, discoverable, platform-specific, and actively maintained
New platforms enter production with complete operational documentation
Engineers spend less time repeatedly explaining known issues
Important operational knowledge no longer depends on individual engineers or undocumented conversations
Where this role can lead
This position sits within the Infrastructure L3 Support team and will grow alongside the function. You will initially establish and own its knowledge system while building deeper expertise in our infrastructure and GPU platforms. For the right person, there is a natural development path toward an L3 Support Engineer position as the team and your platform knowledge mature. We develop people as deliberately as we develop our infrastructure and documentation