NVIDIA Reference Architectures: Deep Dive#

The following sections show how NVIDIA Reference Architecture guidance translates into specific Enterprise AI Factory designs. Each deep dive summarizes the reference configuration, scaling model, networking topology, and workload fit for one AI Factory family: NVIDIA RTX PRO AI Factory, NVIDIA HGX AI Factory, and NVIDIA NVL72 AI Factory.

Together, these examples show how a common RA discipline can support different deployment needs, from flexible PCIe-based infrastructure in existing data centers, to dense HGX clusters for demanding AI and HPC workloads, to rack-scale NVL72 systems for frontier-scale AI.

Reference Architecture for NVIDIA RTX PRO AI Factory#

NVIDIA RTX PRO AI Factory reference configurations are designed for universal enterprise AI workloads, including generative and agentic AI inference, fine-tuning, perception AI, data analytics, HPC, and visual computing. These configurations use PCIe-based NVIDIA-Certified compute nodes with NVIDIA RTX PRO 6000 or RTX PRO 4500 Blackwell Server Edition GPUs, and scale from a minimum of 4 nodes in 4-node increments up to 32 nodes.

2-8-5-200 Reference Configuration is the foundational RTX PRO AI Factory node design. It supports 2 CPUs balanced with 8 GPUs, 5 network adapters, and 200 GbE of east-west bandwidth per GPU. This configuration includes a full east-west fabric for distributed workloads and north-south networking for storage, uplink, and management traffic.

  • Eligible GPU: NVIDIA RTX PRO™ 6000 or RTX PRO™ 4500 Blackwell Server Edition; H200 NVL optional in select configurations

  • Eligible CPU: AMD EPYC processors: Milan, Genoa, Turin; Intel Xeon Scalable processors: Sapphire Rapids, Emerald Rapids, Granite Rapids

  • East-west networking: 4x NVIDIA BlueField-3 B3140H

  • North-south networking: 1x NVIDIA BlueField-3 B3220

  • Ideal Workloads: Distributed inference, fine-tuning, perception AI, HPC, data analytics, and visual computing

_images/white-paper-rtx-pro-2-8-5-200-config.png

Figure 1 NVIDIA RTX PRO AI Factory 2-8-5-200 reference configuration shows the foundational PCIe-optimized node design with (8) RTX PRO GPUs, full east-west fabric, and north-south connectivity for storage, uplink, and management traffic.#

Volume 2-4-3-200 and 2-2-3-400 Reference Configurations provide lower-cost RTX PRO AI Factory options for use cases that need east-west networking but do not require the full compute of the foundational configuration. These patterns support 2 CPUs with 2 or 4 GPUs, 3 network adapters, and 200 GbE of east-west bandwidth per GPU. They preserve east-west connectivity while reducing adapter count, cost, power, cabling, and form-factor requirements.

  • Eligible GPU: NVIDIA RTX PRO 6000 or RTX PRO 4500 Blackwell Server Edition; H200 NVL optional in select configurations

  • Eligible CPU: AMD EPYC processors: Milan, Genoa, Turin; Intel Xeon Scalable processors: Sapphire Rapids, Emerald Rapids, Granite Rapids; NVIDIA Vera CPU eligible in select configurations

  • East-west networking: 2x NVIDIA BlueField-3 B3140H

  • North-south networking: 1x NVIDIA BlueField-3 B3220

  • Ideal Workloads: Inference, fine-tuning, perception AI, data analytics, visual computing, and workloads that benefit from east-west connectivity but do not require maximum per-node fabric density

_images/white-paper-rtx-pro-volume-config.png

Figure 2 NVIDIA RTX PRO AI Factory volume reference configurations show lower-cost 2-GPU and 4-GPU node designs that preserve east-west connectivity while reducing adapter count, cabling, power, and form-factor requirements.#

Extreme Volume 2-2-1-NA, 2-4-1-NA, and 2-8-1-NA Reference Configurations are north-south-only RTX PRO AI Factory patterns for cost-optimized deployments. These configurations support 2 CPUs with 2, 4, or 8 GPUs and 1 network adapter for north-south traffic. Because east-west networking is not included, the bandwidth field is not applicable. These designs are intended for workloads that fit within a single GPU or do not require GPU-to-GPU communication across nodes.

  • Eligible GPU: NVIDIA RTX PRO 6000 or RTX PRO 4500 Blackwell Server Edition; H200 NVL optional in select configurations

  • Eligible CPU: AMD EPYC processors: Milan, Genoa, Turin; Intel Xeon Scalable processors: Sapphire Rapids, Emerald Rapids, Granite Rapids; NVIDIA Vera CPU eligible in select reference configurations

  • East-west networking: Not applicable

  • North-south networking: 1x NVIDIA BlueField-3 B3220

  • Ideal Workloads: Generative and agentic AI inference, visual computing, perception AI, and data analytics where workloads do not need to span GPUs across nodes

_images/white-paper-rtx-pro-extreme-volume-config.png

Figure 3 NVIDIA RTX PRO AI Factory extreme volume reference configurations show north-south-only node designs for inference, visual computing, perception AI, and data analytics workloads run within a GPU.#

Scaling NVIDIA RTX PRO AI Factory Clusters#

NVIDIA RTX PRO AI Factory reference configurations scale from one scalable unit to eight scalable units, or from 4 nodes to 32 nodes. In the 2-8-5-200 foundational configuration, this represents a scale range from 32 GPUs to 256 GPUs. NVIDIA publishes guidance for each step in that range, so customers can grow capacity without designing every intermediate cluster size from scratch.

As RTX PRO clusters scale, the network topology evolves with the deployment size. For smaller deployments up to 4 scalable units, or 16 nodes, NVIDIA RAs use a consolidated network design in which north-south and east-west traffic share the same fabric. This provides a simpler topology for smaller clusters while still supporting the required traffic patterns.

For larger deployments above 16 nodes, NVIDIA RAs move to dedicated north-south and east-west networks. This separation gives GPU compute traffic, customer uplink, storage, and management services clearer paths as the cluster grows, helping preserve performance, supportability, and operational predictability.

Note

NVIDIA Spectrum™-4 SN5600 switches are depicted in both the NVIDIA RTX PRO and HGX AI Factory scaled topologies. NVIDIA Spectrum-4 SN5600 and SN5610 switches are interchangeable in these configurations.

_images/white-paper-rtx-pro-scaling-topology.png

Figure 4 NVIDIA RTX PRO AI Factory with rail-optimized topologies scales from 4 to 32 nodes, or 32 to 256 GPUs. Up to 16 nodes, NVIDIA RAs use a consolidated north-south and east-west network fabric. Above 16 nodes, the design transitions to dedicated north-south and east-west networks.#

Reference Architecture for NVIDIA HGX AI Factory#

NVIDIA HGX AI Factory reference configurations are designed for enterprise AI power users that need dense compute, large GPU memory, and high-speed interconnect for demanding AI and HPC workloads. These configurations use NVIDIA-Certified 8-GPU HGX servers with NVIDIA HGX B300, combining scale-up GPU architecture inside the node with high-bandwidth scale-out networking across nodes.

The foundational HGX AI Factory configuration is built for workloads that require more than flexible PCIe acceleration. It is designed for large language model training and fine-tuning, large-scale distributed inference, GPU-accelerated data analytics, and scientific research. With full east-west networking between nodes and north-south networking for storage, uplink, and management traffic, the design helps enterprises scale performance without introducing deliberate bottlenecks into the cluster.

2-8-9-800 Reference Configuration is the foundational NVIDIA HGX AI Factory node design. It supports 2 CPUs balanced with 8 NVIDIA HGX B300 GPUs, 9 network adapters, and 800 GbE of average east-west bandwidth per GPU. This pattern provides a fully networked HGX design for distributed training, fine-tuning, inference, data analytics, and scientific computing workloads, and scales from 4 nodes in 4-node increments up to 128 nodes.

  • Eligible GPU: NVIDIA HGX™ B300

  • Eligible CPU: AMD EPYC processors: Milan, Genoa, Turin; Intel Xeon Scalable processors: Sapphire Rapids, Emerald Rapids, Granite Rapids

  • East-west networking: 8x NVIDIA ConnectX-8 SuperNICs

  • North-south networking: 1x NVIDIA BlueField-3 B3220

  • Ideal Workloads: AI training, fine-tuning, large-scale distributed inference, data analytics, and scientific research

Compared with the NVIDIA RTX PRO AI Factory foundational configuration, the HGX configuration keeps the same 2-CPU, 8-GPU node shape but increases the network density from 5 adapters to 9 and raises east-west bandwidth from 200 GbE per GPU to 800 GbE per GPU. That fourfold increase in east-west bandwidth is what makes distributed training and other communication-heavy workloads practical at cluster scale.

_images/white-paper-hgx-2-8-9-800-config.png

Figure 5 NVIDIA HGX AI Factory 2-8-9-800 reference configuration shows an 8-GPU scale-up node with balanced CPU, GPU, and networking resources for demanding multi-GPU workloads.#

Scaling NVIDIA HGX AI Factory Clusters#

NVIDIA HGX AI Factory reference configurations scale from one scalable unit to 32 scalable units, or from 4 nodes to 128 nodes. In the 2-8-9-800 foundational configuration, this represents a scale range from 32 GPUs to 1,024 GPUs. NVIDIA publishes guidance across this range so customers can grow capacity without designing each intermediate cluster size from scratch.

As HGX clusters scale, the network topology becomes more specialized. Smaller configurations may use a consolidated fabric, but at 16 nodes NVIDIA RAs recommend moving to dedicated north-south and east-west networks. This transition separates GPU compute traffic from customer uplink, storage, management, and support services, helping preserve performance and operational predictability as the cluster grows. Larger design points (e.g. 32 SU) rely on a spine-leaf network architecture for both the north-south and east-west networks.

Once the networks are dedicated, NVIDIA RAs provide both single-plane and dual-plane east-west options. A single-plane design offers a simpler deployment model, while a dual-plane design provides additional resilience and path diversity. Both are published design patterns, allowing customers and partners to choose the topology that fits their scale, availability needs, and operating model.

_images/white-paper-hgx-scaling-topology.png

Figure 6 NVIDIA HGX AI Factory with rail-optimized topologies scales from 4 to 128 nodes, or 32 to 1,024 GPUs. At 16 nodes, NVIDIA RAs recommend transitioning to dedicated north-south and east-west networks. Dedicated east-west fabrics can be deployed as single-plane or dual-plane designs.#

Reference Architecture for NVIDIA NVL72 AI Factory#

NVIDIA NVL72 AI Factory reference configurations are designed for frontier-scale AI and HPC workloads that require rack-scale compute, high-speed GPU interconnect, and advanced thermal design. Unlike RTX PRO and HGX configurations, which are described primarily as node-level building blocks, NVIDIA NVL72 is a rack-scale architecture. The tray is described by the C-G-N-B notation, but the scalable unit is the full rack.

The NVIDIA NVL72 AI Factory is built on NVIDIA GB300 NVL72 and is designed for large-scale AI training, fine-tuning, inference, scientific research, and data analytics. Fifth-generation NVIDIA NVLink, FP4 Tensor Cores, and advanced thermal management enable trillion-parameter model training and massive mixture-of-experts models. This is the configuration for workloads that need rack-scale infrastructure operating as a single compute domain.

NVL72 2-4-5-800 Reference Configuration is the foundational NVIDIA NVL72 AI Factory tray design. Each tray supports 2 NVIDIA Grace CPUs, 4 NVIDIA Blackwell GPUs, 5 network adapters, and 800 GbE of east-west bandwidth per GPU. The networking design includes 4 NVIDIA ConnectX-8 adapters for east-west traffic and 1 NVIDIA BlueField-3 B3240 DPU for north-south traffic. At the rack level, the configuration includes 18 trays, 36 NVIDIA Grace CPUs, and 72 NVIDIA Blackwell Ultra GPUs.

  • Eligible GPU: NVIDIA Blackwell Ultra

  • Eligible CPU: NVIDIA Grace CPU

  • East-west networking: 4x NVIDIA ConnectX-8 adapters

  • North-south networking: 1x NVIDIA BlueField-3 B3240

  • Scaling unit: 1 rack, with 18 trays/nodes

  • Scale range: 1 to 8 racks, or 72 to 576 GPUs

  • Best fit: Frontier-scale AI training, fine-tuning, inference, scientific research, and data analytics

NVL72 changes the deployment conversation because customers scale by rack, not by four-node increments. Deployments begin with one rack and expand in single-rack increments. This makes power, cooling, and data center readiness prerequisite design considerations. For customers with the workload profile and facility capacity to support it, NVIDIA NVL72 AI Factory provides the rack-scale architecture needed for the largest and most compute-intensive AI workloads.

_images/white-paper-nvl72-2-4-5-800-config.png

Figure 7 NVIDIA NVL72 AI Factory 2-4-5-800 reference configuration is built on NVIDIA GB300 NVL72 and scales from 1 to 8 racks, or 72 to 576 GPUs. Each rack includes 18 trays with 36 NVIDIA Grace CPUs and 72 NVIDIA Blackwell GPUs. The configuration provides 800 GbE of east-west bandwidth per GPU and is designed for frontier-scale AI training, fine-tuning, inference, scientific research, and data analytics.#

Scaling NVIDIA NVL72 AI Factory Clusters#

NVIDIA NVL72 AI Factory reference configurations scale from 1 rack to 8 racks, or from 72 GPUs to 576 GPUs. Unlike RTX PRO and HGX, where the scalable unit is four nodes, the NVL72 scalable unit is one full rack with 18 trays. Customers expand by adding racks, not individual four-node blocks.

As NVL72 clusters scale, the network architecture remains deliberately consistent. NVIDIA RAs use dedicated north-south and east-west networks with dual-plane east-west fabrics across the published range. There is no consolidated-network option and no single-plane variant in the NVL72 guidance, because rack-scale deployments are built for workloads that require the GPUs to operate as a large, highly connected compute domain.

This uniform topology matters because constrained networking would undermine the reason customers choose NVL72. At this scale, the network must not become the bottleneck for frontier-scale training, fine-tuning, inference, scientific research, or data analytics. Rather than asking customers to choose among topology trade-offs as they grow, the NVIDIA RA provides one repeatable design pattern that scales by replicating a known-good rack-scale architecture.

_images/white-paper-nvl72-scaling-topology.png

Figure 8 NVIDIA NVL72 AI Factory with rail-optimized topologies scales from 1 to 8 racks, or 72 to 576 GPUs. Across the full range, NVIDIA RAs use dedicated north-south networking and dual-plane east-west fabrics.#

Storage#

As enterprises build AI factories the importance of this data cannot be overstated: access to high quality data directly impacts the performance and reliability of AI models. Data is essential for developing and optimizing AI applications, and it must be fed across all stages of the AI pipeline, from model building to training, tuning, and inference, with varied storage requirements at each stage. It’s fuel for the AI factory.

The NVIDIA-Certified™ Storage program is a comprehensive validation framework designed to help leading storage vendors deliver high-performance storage solutions. It enables partners and customers to deploy to build AI factories that efficiently leverage massive amounts of data for faster, more accurate, and reliable AI models. NVIDIA Enterprise RAs have designated network end points to attach NVIDIA-Certified storage solutions.

NVIDIA-Certified Storage evaluates storage through a broader AI lens. It examines how a platform behaves across different workload patterns, how it responds under stress, and how well it works with NVIDIA-Certified systems, NVIDIA networking, and the operational patterns customers use every day.

By rigorously testing partner solutions against real-world AI workloads and synthetic benchmarks at scale, NVIDIA-Certified Storage helps vendors demonstrate that their technology is ready for these demands. The program evaluates technical merit, validates reference configurations, and connects storage performance to full-stack AI deployment patterns.

The NVIDIA-Certified Storage program has several types of certifications. These storage certifications are matched with corresponding NVIDIA Enterprise RAs at different scales to ensure the storage systems have the performance to support the performance, security, and scale required for large-scale production AI workloads.

Software#

NVIDIA Enterprise RAs include companion guides for deploying the software stack on top of the validated cluster. This guidance connects the bare-metal infrastructure foundation with orchestration, NVIDIA AI Enterprise software, observability, workload sizing, and operational practices so partners and customers can move from cluster design to a production AI Factory environment.

The RAs can be used with industry-standard orchestration tools, including Kubernetes and Slurm, and with NVIDIA AI Enterprise components for enterprise-grade AI software and models. As the program expands, NVIDIA Enterprise RAs will be accompanied by infrastructure ISV software designs (like the Palantir Sovereign AI Operating System Reference Architecture with NVIDIA) that have been tested and validated on NVIDIA RAs, helping partners and customers simplify AI adoption and optimize performance with leading software offerings.