Components#

This section describes the hardware components of the RA.

NVIDIA HGX H100/H200 8-GPU Baseboard#

The HGX 8-GPU H100/H200 baseboard (Figure 1) is an AI powerhouse that enables enterprises to expand the frontiers of business innovation and optimization.

_images/hgx-ai-factory-h100-h200-b200-01.png

Figure 1 HGX 8-GPU H100/H200 baseboard.#

The HGX H100/H200 baseboard combines H100/H200 Tensor Core GPUs with high-speed interconnects to form the world’s most powerful systems. With eight H100/H200 GPUs, the H100/H200 baseboard has up to 640 GB (1128GB for H200) of GPU memory for unprecedented acceleration.

NVIDIA HGX B200 8-GPU Baseboard#

The HGX B200 baseboard (Figure 2) is a Blackwell x86 platform, based on eight B200 GPUs, that delivers 144 petaFLOPs of AI performance. The HGX B200 baseboard delivers the best performance (15 times more than the HGX H100 baseboard) and TCO (12 times more than the HGX H100 baseboard) for x86 scale-up platforms and infrastructure. Each GPU is configurable up to 1 kW per GPU.

_images/hgx-ai-factory-h100-h200-b200-02.png

Figure 2 HGX 8-GPU B200 baseboard.#

Table 1: HGX H100, H200 or B200 GPU and per Node specifications.

Specification

NVIDIA H100 SXM

NVIDIA H200 SXM

NVIDIA B200 SXM

Memory per GPU

80GB HBM3

141GB HBM3e

180GB HBM3e

Memory per 8-GPU Node

640GB HBM3

1.1TB HBM3e

1.44TB HBM3e

GPU Bandwidth

3.35TB/s

4.8TB/s

Up to 8TB/s

GPU Aggregate Bandwidth per 8-GPU Node

26.8TB/s

38.4TB/s

Up to 64TB/s

The HGX H100/H200 baseboard provides 32 petaflops of processing power, making it a leading accelerated platform for AI and HPC. It offers advanced networking up to 400 Gbps with BlueField-3 SuperNICs, enabling AI cloud networking, composable storage, zero-trust security, and GPU compute elasticity in AI designs.

Meanwhile, the HGX B200 delivers 144 petaflops, also positioning it as a top accelerated platform for AI and HPC. It features similar advanced networking capabilities at speeds of up to 400 Gbps with BlueField-3 SuperNICs, supporting AI cloud networking, composable storage, zero-trust security, and GPU compute elasticity in AI designs.

HGX H100, H200 or B200 8-GPU Systems#

NVIDIA-Certified HGX H100, H200 or B200 8-GPU systems are based on a common system design with flexibility for optimizing the configuration to match the cluster requirements. Systems are available in 4-GPU and 8-GPU configurations. The RA was built using an 8 GPU design. An example of this design is shown in Figure 3. Please note, 4 GPU designs can also be used.

_images/hgx-ai-factory-h100-h200-b200-03.png

Figure 3 Example of a HGX H100, H200 or B200 8 GPU system configuration.#

Elements for the HGX H100, H200 or B200 8-GPU NVIDIA-Certified system are listed in Table 2.

Table 2: HGX H100, H200 or B200 8-GPU system elements.

Parameter

Requirement

Target workloads

Large Language Model; Traditional DL Inference Models, HPC

GPU configuration

Eight H100, H200 or B200 SXM GPUs on a H100, H200 or B200 baseboard
• H100 variant has up to 640 GB of GPU memory
• H200 variant has up to 1128 GB of GPU memory
• B200 variant has up to 1440 GB of GPU memory
See the topology diagrams for details.

NVIDIA® NVLink™ and NVSwitch™

HGX H100 and H200 8-GPU Baseboards use a combination of third generation NVSwitch and fourth generation NVLink
• Total Aggregate Bandwidth 7.2TB/s
• GPU-to-GPU Bandwidth 900GB/s

HGX B200 8-GPU Baseboard use a combination of fourth generation NVSwitch and fifth generation NVLink
• Total Aggregate Bandwidth 14.4TB/s
• GPU-to-GPU Bandwidth 1800GB/s

CPU

Intel Emerald Rapids, or Sapphire Rapids
AMD Turin and Genoa

CPU sockets

Two CPU sockets minimum

CPU speed

2.1 GHz minimum base CPU clock

CPU cores

Minimum of 48 physical CPU cores per socket
Recommendation of 56 physical CPU cores per socket

System memory (total across all CPU sockets)

Minimum of 1.5TB system memory
Minimum of 500GB/s memory bandwidth
For optimal performance, system memory should be evenly distributed across all CPU sockets and memory channels and should be fully populated and symmetrically placed on all CPU memory controller (MC) channels.

DPU (North-South)

One NVIDIA® BlueField®-3 DPU per server

PCI Express

Eight Gen5 x16 links and one Gen4 x2 link per HGX H100, H200 or B200 8-GPU baseboard
One Gen5 x16 link per DPU, SuperNIC or adapter

PCIe topology

Balanced PCIe topology with connectivity spread evenly across CPU sockets and PCIe root ports.

Network Adapters/NICs speed (East-West)

Eight NVIDIA® BlueField®-3 SuperNICs per server
Up to 400 Gbps per adapter

Local storage

Local storage recommendations are as follows:
Inference Servers: Minimum 1 TB NVMe drive per CPU socket
Training / DL Servers: Minimum 2 TB NVMe drive per CPU socket
HPC Servers: Minimum 1 TB NVMe drive per CPU socket
1 TB NVMe boot drive

Remote systems management

SMBPBI over SMBus (OOB) protocol to BMC
PLDM T5-enabled
SPDM-enabled

Security

TPM 2.0 module (secure boot)

Control Plane/Management Nodes#

In addition to the standard compute nodes, control plane nodes are required to support the management and access for the customer to the training cluster. The components of the control plane nodes will depend on the software stack used for provisioning and managing the cluster. For example, a configuration using NVIDIA Base Command Manager, Slurm, and Kubernetes together can include seven control plane nodes in total: two for Base Command Manager (with high availability configured), two for Slurm head nodes, and three for Kubernetes control plane nodes. This uses seven of the eight available control plane nodes possible.

Table 3 shows the recommended configuration for the control plane nodes.

Table 3: Control plane node components.

Component

Quantity

Description

CPU

2

• 32C Intel Xeon Gold 6448Y or equivalent
• 32C AMD EPYC 9354

North-South (DPU)

1

NVIDIA BlueField-3 B3220 DPU with two 200G ports and 1Gb RJ45 management port
Other variants can be supported as per Compute Node alternatives

System Memory

Minimum of 256 GB DDR5

Boot Drive

1

1 TB NVMe SSD

Local storage

1

4 TB NVMe SSD. More may be required if image storage is required

BMC

1

1 Gb RJ45 management port

In cases where the existing control plane nodes are missing, the Enterprise RA recommends deploying one set per cluster. Configure these nodes for high availability to keep your cluster running smoothly.

HGX H100, H200 or B200 8-GPU System Networking#

The NVIDIA HGX platform networking configuration enables the highest AI performance and scale, while ensuring cloud manageability and security. It leverages the NVIDIA expertise in AI cloud data centers and optimizes network traffic flow:

  • East/West (Compute Network) traffic: This refers to traffic between NVIDIA HGX systems within the cluster, typically for multi-node AI training, AI fine tuning, HPC collective operations, and other workloads.

  • North/South (Customer and Storage Network) traffic: This involves traffic between NVIDIA HGX systems and any external resources including cloud management and orchestration systems, remote data storage nodes, and other parts of the data center or the Internet.

Combined with NVIDIA Spectrum-X Ethernet, NVIDIA HGX H100, H200 or B200 8-GPU platform delivers the highest performance for DL training and inference, data science, scientific simulation, and other modern workloads. The following sections describe the recommended NVIDIA HGX platform configurations with their associated NVIDIA networking platforms.

Compute (Node East/West) Ethernet Networking#

To deliver the highest AI performance, NVIDIA recommends the BlueField-3 (BF-3) SuperNIC smart network adapters.

The BlueField-3 SuperNIC offers up to 400 Gb/s, low-latency network connectivity between GPUs in the AI cluster, featuring RDMA and RoCE acceleration, with NVIDIA® GPUDirect® and GPUDirect Storage technologies. For data centers that deploy Ethernet, BF-3 offers a range of advancements including RoCE optimizations and multi-tenancy, which are key to in AI cloud environments.

BF-3 is a central part of the NVIDIA Spectrum-X networking platform, which also features NVIDIA Spectrum-4 switches. At its core, the BlueField-3 SuperNIC emerges as a novel network accelerator, purpose-built to supercharge hyperscale AI workloads. It ensures lightning-fast and efficient communication between GPU servers, for peak AI workload efficiency. Furthermore, the SuperNIC can establish a secure multi-tenant data center environment, while guaranteeing deterministic and isolated performance for tenant jobs. With its power-efficient HHHL PCIe design, the BlueField-3 SuperNIC seamlessly integrates into NVIDIA HGX systems, significantly boosting performance for AI traffic on the east-west network inside the clusters.

Multi-node deployments with an NVIDIA HGX H100, H200 or B200 8-GPU platform should adhere to the following total compute network bandwidth per GPU recommendations.

The BF-3 SuperNIC supports two operation modes (NIC and DPU modes), with NIC mode as the default. If the SuperNIC is set to DPU mode, the DPU 1 GbE out-of-band management port must be connected.

The NVIDIA multi-node software stack deployment is optimized for several GPU to NIC ratios. Partners are recommended to accommodate NICs for the 1:1 GPU to NIC ratio.

Total Minimum Compute Network Bandwidth

  • >200 GB/s (8x 200 Gb/s or 4x 400 Gb/s NICs)

Total Recommended Compute Network Bandwidth

  • 400 GB/s (8x 400 Gb/s NICs)

For the NICs used in HGX H100, H200 or B200 8-GPU systems, NVIDIA recommends the following NVIDIA products for the Compute (East-West) Networking.

Table 4: HGX H100, H200 or B200 8-GPU recommended SuperNICs for East-West Network.

Product

PCIe Card Form Factor

Connector Required

Applicable Topologies

NVIDIA BlueField-3 B3140H E-series HHHL DPU, 400GbE (default mode) /NDR IB, Single-port QSFP112

HHHL
PCIe Gen5 x16

75W system power supply through the PCIe x16 interface
1G/USB Management Interfaces
QSFP112 connector cages

Switch in Virtual or Synthetic Mode / Switch in Base Mode

NVIDIA BlueField-3 B3220L E-Series FHHL DPU, 200GbE (default mode) /NDR200 IB, Dual-port QSFP112

FHHL
PCIe Gen5 x16

75W system power supply through the PCIe x16 interface
1G/USB Management Interfaces
QSFP112 connector cages

Switch in Virtual or Synthetic Mode / Switch in Base Mode

Converged (Node North/South) Ethernet Networking#

This section describes the BlueField-3 role for the North-South (N-S) Ethernet network in the NVIDIA HGX platform and the recommended BlueField-3 models for this infrastructure.

The NVIDIA BlueField-3 data processing unit (DPU) is a 400 Gb/s infrastructure compute platform that enables organizations to securely deploy and operate NVIDIA HGX AI data centers at massive scales. BF-3 DPU is optimized for the N-S network and the BF-3 SuperNIC is optimized for the E-W, AI compute fabric.

BlueField-3 offers several essential capabilities and benefits within the NVIDIA HGX platform:

  • Workload Orchestration: NVIDIA BlueField-3 serves as an optimized compute platform for the data center control-plane, enabling automated provisioning and elasticity. This empowers NVIDIA HGX AI cloud platforms to scale resources dynamically based on fluctuating demand, ensuring efficient allocation of computing resources for transient AI workloads.

  • Storage Acceleration: BlueField-3 provides advanced storage acceleration features that optimize data storage access. Its innovative BlueField SNAP technology enables remote storage devices to function as local, improving AI performance and streamlining cloud operations.

  • Secure Infrastructure: BlueField-3 operates in a highly secure zero-trust mode and functions independently from the host, significantly enhancing the security of the NVIDIA HGX platform. In addition, BlueField-3 enables a wide range of accelerated security services, including next-generation firewall and micro-segmentation, bolstering the overall security posture of the infrastructure.

The NVIDIA BlueField-3 DPU integration enhances NVIDIA HGX platforms by improving resource management, performance, security, and scalability, making them well-suited for AI workloads in datacenter environments.

NVIDIA BlueField-3 DPUs support multiple operating modes, including Embedded Function (ECPF) or DPU mode, which is commonly used as a default. In this mode, the Arm subsystem on the DPU owns and manages NIC resources, and network traffic typically passes through a virtual switch on the DPU before reaching the host, adding an extra layer of control.

To securely control and manage the NVIDIA HGX platform through the BlueField DPU, NVIDIA has integrated an onboard BMC into its BlueField products. The onboard BlueField BMC enables the provisioning and management of the BlueField and NVIDIA HGX platforms using standard tools, including Redfish APIs. It includes an external root-of-trust to ensure that the BMC firmware is secured. The main interface of the BlueField BMC is a 1GbE out-of-band management port that connects to the data center and management network.

To ensure platform compatibility with the recommended BlueField DPUs, NVIDIA is also providing form factor and power guidance. Some BlueField-3 configurations for the N/S network has greater than 75 W power consumption and requires an external PCIe power connector capable of 75 W minimum.

For the BlueField-3 DPUs used in the N/S network for NVIDIA HGX platforms, NVIDIA recommends the following BlueField-3 products listed in Table 5.

Table 5: HGX H100, H200 or B200 8-GPU recommended BF-3 DPUs for North-South Network.

Product

PCIe Card Form Factor

Connector Required

Applicable Topologies

NVIDIA BlueField-3 B3220 P-series FHHL DPU, 200GbE (default mode) /NDR200 IB, Dual-port QSFP112

FHHL
PCIe Gen5 x16 with x16 PCIe extension option

8-pin ATX 12V PCIe
1G/USB Management Interfaces
QSFP112 connector cages

All

NVIDIA BlueField-3 B3240 P-Series FHHL DPU, 400GbE (default mode) /NDR IB, Dual-port QSFP112

FHHL
PCIe Gen5 x16 with x16 PCIe extension option

8-pin ATX 12V PCIe
1G/USB Management Interfaces
QSFP112 connector cages

All