Overview of the Reference Architecture (RA)#
The NVIDIA RA using the 2-8-9-400 node architecture with NVIDIA HGX H100/H200/B200 8-GPU and NVIDIA Spectrum-X Networking is optimized for workloads of multi-node AI or hybrid applications. The RA is a modular architecture based on NVIDIA-Certified HGX H100, H200, or B200 systems, each with eight H100, H200, or B200 SXM GPUs. Using a four-node scalable unit (SU) (refer to Figure 10), this RA scales up to 128 NVIDIA-Certified HGX H100, H200, or B200 systems for a total of 1,024 SXM GPUs.
A fully tested system scales to thirty-two SUs (Scalable Units). Larger clusters can be built based on customer requirements.
Flexible rail-optimized end-of-row network architecture that can accommodate modifications in the rack layout and number of servers per rack.
Hardware support is available through the fulfillment OEM and channel partners. Software support from NVIDIA is based on a per-GPU paid subscription of NVIDIA AI Enterprise
The NVIDIA HGX H100/H200/B200 and NVIDIA Spectrum-X Networking Platform RA enables the following use cases:
AI Inference—Large (per node) and Medium (per GPU) model parameter inference workload.
AI Training––Large to small model training and fine tuning based on cluster sizing.
Our Sizing Guides provide guidance on typical characteristics and models supported for various workloads.
For all use cases, this architecture is ideal for multi-user, single-tenant workloads. Specifically, the logical design and software is streamlined for deployment and maintenance ease by tailoring the configuration to one where users are all part of the same enterprise, and accounting and access control can be consolidated.
Similarly, Kubernetes is the modern foundation of enterprise AI work, and this RA is architected for deploying Kubernetes and Kubernetes-dependent applications and tooling.