Define the hardware architecture of our AI accelerator systems, spanning server board design, rack integration, and datacenter-level deployment
Translate system architecture requirements for compute, memory, and interconnect/fabric, defined in collaboration with the broader system architecture team and product division, into concrete hardware specifications
Specify physical interconnect and fabric infrastructure, defining the architectural requirements and trade-offs that the AI Infrastructure Systems division executes against, including how interconnect and fabric characteristics impact system-level workload performance across scale-up and scale-out topologies
Define host, storage, and networking integration and organization at the hardware level, covering PCIe and CXL interconnect standards, as well as system-level power budgeting
Ensure scale-up and scale-out designs sustain target performance as datacenter systems grow from single nodes to large clusters
Define operability and RAS (reliability, availability, serviceability) requirements across redundancy architecture, hot-swap capability, telemetry and management interfaces (BMC/IPMI/Redfish), and fault containment to ensure datacenter platforms are manageable and dependable in production
Take a leading technical role in the system architecture team, interfacing directly with key partners and internal stakeholders to align on architectural decisions on the datacenter hardware. This includes interactions with the AI Infrastructure Systems division, silicon, product management, software, and customers, system integrators, and ecosystem partners on architecture requirements — with commercial engagement and design ownership sitting with the AIIS Director
Drive methodology and best practices for datacenter system architecture as the team scales