Introducing NVIDIA Reference Architectures#

NVIDIA Reference Architectures provide clear, consolidated cluster design guidance for NVIDIA global system partners and enterprise customers building Enterprise AI Factories. They span compute, networking, storage, and infrastructure software, bringing NVIDIA-Certified Systems, NVIDIA-Certified Storage, NVIDIA Networking, and NVIDIA AI Enterprise software into integrated designs for production AI infrastructure.

These reference architectures are based on decades of NVIDIA experience in accelerated computing, including lessons learned from hyperscale cloud, supercomputing, enterprise AI deployments, and NVIDIA’s own internal AI factory. NVIDIA architects develop, test, deploy, and operate AI infrastructure at scale, then package those learnings into recommended design patterns that partners can build from instead of starting from scratch.

That shared learning helps reduce guesswork and deployment risk. Small infrastructure choices can become major issues at scale, such as cabling and transceiver decisions that affect thermals, delivery time, and customer experience. System balance matters as well: CPU, GPU, storage, and networking ratios may look acceptable for a few nodes or a single workload, but imbalanced designs can limit distributed inference, reduce utilization, or create operational bottlenecks as clusters scale.

NVIDIA RAs turn these lessons into published guidance for well-balanced, supportable systems in which bottlenecks caused by individual components are minimized. The guidance helps partners design flexible, cost-effective configurations that improve performance, utilization, uptime, total cost of ownership, and supportability.

Bringing Reference Architectures to Market#

NVIDIA brings reference architectures to market through a repeatable develop, establish, endorse, and scale model.

  • Develop: NVIDIA builds, tests, and operates AI infrastructure internally, using DGX systems and NVIDIA’s own AI factory experience to validate system, cluster, and software deployment patterns.

  • Establish: Those learnings are documented across three levels: OEM system-level certification, cluster-level NVIDIA Reference Architecture guidance, and AI ecosystem software reference designs. NVIDIA-Certified Systems cover the server or node. NVIDIA Reference Architectures define how those systems come together as a cluster. NVIDIA software reference designs and validated designs help customers understand what to run on top of the infrastructure.

  • Endorse: To make partner claims verifiable, NVIDIA reviews qualifying OEM cluster designs through the NVIDIA Design Review Board (DRB), a technical review process led by NVIDIA engineers. The DRB evaluates whether the design follows applicable NVIDIA RA patterns and can be deployed reliably as an Enterprise AI Factory solution.

  • Scale: Global system partners package endorsed architectural guidance into differentiated Enterprise AI Factory offerings. They build on NVIDIA Reference Architectures, then add infrastructure choices, services, software integrations, deployment expertise, and industry-specific solutions.

DRB endorsement gives customers a published signal that NVIDIA reviewed the design for architectural fidelity, balance, and ease of deployment. It also gives partners a repeatable foundation they can productize, enable across sales teams, and scale across customers.

The endorsement process has three steps: the partner designs an OEM Reference Architecture using NVIDIA-Certified nodes and NVIDIA RA patterns; NVIDIA engineers review the design through the DRB; and designs that pass are recognized and published as NVIDIA-endorsed designs. Endorsed designs are published on NVIDIA web and documentation properties so customers can check which OEM designs have passed review, which configuration each design uses, the supported cluster size range, and which endorsement categories the design carries.

Endorsement category

Role in the DRB process

Infrastructure Configuration

Required baseline endorsement. Confirms that the system and cluster design follows the applicable NVIDIA RA infrastructure pattern.

NVIDIA Spectrum-X Compatible

Recommended additional endorsement. Confirms that the Spectrum-X fabric design aligns to the requirements of AI east-west networking.

Networking Logical Architecture

Recommended additional endorsement. Confirms the reviewed logical topology for north-south and east-west networking at the applicable scale.

Table 1: This table represents the currently available NVIDIA Design Review Board endorsements for partner AI Factory designs.

Enterprise AI Factory Families#

NVIDIA Reference Architectures support three primary Enterprise AI Factory configurations available from global system partners. These configurations provide optimized starting points for different workload requirements, deployment scales, and real-world enterprise data center constraints including space, power, cooling, and integration with existing infrastructure.

All three support a broad range of enterprise AI workloads, but each is optimized for a different customer need. Actual enterprise deployments may combine multiple configurations to address different workloads and requirements.

NVIDIA RTX PRO AI Factory#

The NVIDIA RTX PRO AI Factory, built on NVIDIA RTX PRO Servers with NVIDIA AI software and NVIDIA networking, delivers workload acceleration for data centers with practical space, power, and cooling limits. It is designed for environments that require PCIe connectivity and air cooling, while supporting a broad range of enterprise AI workloads. It is especially strong for gen AI inference and visual computing tasks such as graphics rendering in media and entertainment, video search and summarization, and digital twin visualization, making it well suited for enterprises that need flexible AI infrastructure within existing data center constraints.

NVIDIA HGX AI Factory#

The NVIDIA HGX AI Factory, built on NVIDIA HGX B300 with NVIDIA AI software and NVIDIA networking, is designed for AI power users who need dense compute, large GPU memory, and high-speed interconnects. It enables training and fine-tuning of large language models while accelerating inference. The NVIDIA HGX B300 system delivers 15x higher token throughput compared to the NVIDIA HGX H100, driven by Blackwell advancements in FP6 and FP4 precision support. Ultra-fast interconnects reduce data transfer bottlenecks and support large-scale analytics, making HGX well suited for demanding enterprise AI workloads.

NVIDIA NVL72 AI Factory#

The NVIDIA NVL72 AI Factory, built on NVIDIA GB300 NVL72 with NVIDIA AI software and networking, delivers rack-scale performance for AI at massive scale. It offers 50x increased AI Factory output and strong energy efficiency at scale. The technical innovations at the heart of the NVL72 AI Factory, including the fifth-generation NVLink interconnect, FP4 Tensor Cores, and advanced thermal management, enable trillion-parameter AI model training and massive-scale mixture-of-experts models, making it well suited for deployments supporting the largest models and most computationally intensive training and inference workloads.