Summary#
The RA is the leading data center scale architecture to meet the demanding and growing needs of AI training. The RA represents the architecture used by NVIDIA for internal AI model training and HPC research and development. This specific RA focuses on the H100, H200 or B200 baseboard delivered in HGX servers. It is interconnected using Spectrum-X for compute and converged connectivity. This document references other RAs, which provide details on some components described in this RA at a high level.
The RA represents a complete system of hardware and software necessary to implement data centers for high-performance AI functionality. The combination of all these elements keeps systems running reliably, with maximum performance, and enables users to push the bounds of state-of-the-art. The RA is designed to both support the workloads of today and grow to support tomorrow’s applications.