Abstract#

This NVIDIA Reference Architecture (RA) is a practical design guide for an NVIDIA HGX AI Factory based on a 2-8-9-400 infrastructure configuration (2 CPUs, 8 GPUs, 9 NICs at 400 Gb/s bandwidth per GPU) with NVIDIA HGX™ H100, H200, or B200 servers featuring eight GPUs per node, NVIDIA ConnectX™ SuperNICs, NVIDIA BlueField®-3 DPUs, and NVIDIA Spectrum-X™ Ethernet networking at the 32, 64, and 128 node design points.

The NVIDIA HGX AI Factory is designed to support enterprise AI training, fine-tuning, and inference workloads with industry-leading performance in an air-cooled form factor, as well as high-performance computing (HPC) applications.

This document outlines the hardware components that define this scalable and modular architecture, including guidance regarding the scalable unit (SU) design and specifics of Ethernet fabric topologies.