Define the platform-level architecture for our datacenter AI accelerator systems, spanning server board design, rack integration, and datacenter-level deployment
Translate system-level compute, memory, and interconnect requirements — defined in collaboration with the broader System Architecture team and product division — into concrete hardware platform specifications
Specify physical interconnect infrastructure — defining the architectural requirements and trade-offs that the AI Infrastructure Systems division executes against, including how interconnect characteristics impact system-level workload performance
Define host, storage and networking integration and organization at the hardware level, covering PCIe and CXL, as well as system-level power budgeting
Ensure scaled-up and scaled-out designs sustain target performance as systems grow from single nodes to large clusters
Define operability and RAS (reliability, availability, serviceability) requirements across redundancy architecture, hot-swap capability, telemetry and management interfaces (BMC/IPMI/Redfish), and fault containment to ensure platforms are manageable and dependable in production
Take a leading technical role in the system architecture team, interfacing directly with key partners and internal stakeholders to align architecture decisions, including the AI Infrastructure Systems division, silicon, product management, software, and customers, system integrators and ecosystem partners on architecture requirements — with commercial engagement and design ownership sitting with the AIIS Director
Drive methodology and best practices for platform-level design as the team scales