NVIDIA Requirements for AI Clouds v2.2
This page preserves version 2.2. For current requirements, see version 2.4.
Purpose and Intent
This document serves three main purposes:
- Setting requirements for NVIDIA Cloud Partners (NCPs) delivering GPU capacity to NVIDIA
This is the primary requirements document from NVIDIA to any NCP providing NVIDIA GPU/AI compute and software services. These requirements cover the full stack of AI cloud infrastructure services and operations needed to run NVIDIA DGX Cloud, expanding on the NVIDIA hardware reference design.\ - Providing a reference set of requirements for the industry
NVIDIA is publishing this document openly so that NCPs, GPU datacenter operators, and AI practitioners can use it as a reference for the capabilities a large GPU consumer requires\ - Defining NVIDIA’s service delivery expectations
NVIDIA expects services to be delivered as Generally Available to all â not bespoke implementations built for NVIDIA alone. NVIDIA expects operational excellence, transparent communication, and proactive engagement from all partners.
NVIDIA will consider additional services that an NCP offers or plans to offer beyond what is described here.
Service Delivery SLAs
NCPs should be able to demonstrate ability to meet below SLA by category and operational requirements to be considered for offtake.
Services Delivery Timelines
The NCP must demonstrate API readiness, transport establishment at least 12 weeks ahead of GPU delivery, and the ability to provide Dev capacity (CPU nodes only) with the API integrated 6 weeks prior to GPU and cluster delivery.
One key request is for early access to ancillary compute nodes to act as the Data Mover function. This will help us pre-position data into the data center for use when GPUs are available. Access to Data Mover compute (and target storage) should be available ~2 weeks ahead of GPU cluster delivery.
SLA and SLO
Managed K8s
- Control Plane SLA target: Financially-backed 99.95%+ uptime for production.
Storage
- Performance (QoS): Must provision needed throughput requested for minimum bandwidth and IOPS.\
- Home Directory Storage:\
- Availability: Over 99% availability for unplanned incidents. Exclusive of scheduled maintenance.\
- Durability: Over 99.99% for any FS less than 1 PB\
- High Speed Storage Service Requirements:\
- Availability (SLO): Must meet 99.99% availability in a 30-day rolling SLO exclusive of maintenance\
- High-Speed Storage Filesystem Requirements\
- End to End Availability: Over 99.5% uptime per PB\
- Durability: Over 99.999% durability per PB annually
Operational Requirements
- Dedicated Technical specialist/engineer available to NVIDIA\
- Slack channel monitored by technical specialist / engineer\
- 24x7 support available per partner standard incident severity procedures\
- Service impacting incidents, planned, and unplanned maintenance events are communicated to NVIDIA.\
- For planned maintenance, NVIDIA can schedule maintenance windows via APIs / console tools - avoiding unexpected outages + the ability for NVIDIA to provide feedback.\
- NCP to remediate critical vulnerabilities in a timely manner while providing transparent disclosures of any issues
Telemetry Delivery Method
NCP shall deliver all required telemetry, including metrics and logs, in a manner that allows for ingestion into DGX Cloud systems. The preferred methodology is natively via the OpenTelemetry Protocol with a latency of no longer than 120 seconds.
Exemplar Cloud Workload Performance
NVIDIA Exemplar Cloud seeks to improve performance per TCO with hardware and software recipes, references, tools, and capabilities. Run the latest publicly available release from https://github.com/NVIDIA/dgxc-benchmarking (Always pick the latest release version from the GH repo) to be successfully completed on 1 uniform HW cluster type. Please run all the workloads for a given release and share the results in the template below.
Compute and Network Provisioning
This section outlines the requirements for provisioning compute and network. Compute instances can be provided as either Bare Metal instances (via BMaaS) or Virtual Machines (via VMaaS) to support the NVIDIA DGX Cloud engagement. All operations must be controlled via a fully documented and secure API, gRPC or REST preferred. All systems are expected to scale and perform at scale.
General, Compute and Lifecycle Management
Boot Process and Disks
SDN and Virtual Networking
This section covers the virtual networking requirements. Physical transport and network are discussed later in the document.\
Kubernetes As a Service (KaaS) Requirements
Kubernetes Conformance, Versioning, & Compliance
Kubernetes Operational Excellence
Robust K8s Security
Kubernetes Component and Extension Requirements
Kubernetes Functionality
Security and Identity Management
Identity & Access Management (IAM)
Cryptography and Key Management
Network Isolation & Encryption
Edge Network Security
Hardware Security & Compliance
Breakfix Requirements
The NCP must provide a specific “Breakfix API” to support fleet reliability. Any node-level remediation must not impact other parts of the tenancy; specifically, NVLink must be re-configured properly to take a node out of the tenancy.
The API must enable the following actions:
Telemetry Requirements
The telemetry requirements are comprised of two core components that require alignment between DGX Cloud and the NCP:
- Delivery Method: How telemetry will be delivered by NCP to DGX Cloud for ingestion\
- Telemetry Scope: What telemetry the NCP will deliver to DGX Cloud
Delivery Method
NCP shall deliver all required telemetry, including metrics and logs, in a manner that allows for ingestion into DGX Cloud systems. The preferred methodology is natively via the OpenTelemetry Protocol with a latency of no longer than 120 seconds.
Telemetry Scope
DGX Cloud will provide the NCP with a detailed specification document with the required metrics and logs. Upon receipt, the NCP shall be required to provide a formal written response detailing the following:
- Confirmation of its ability to deliver the specified metrics and logs.\
- Projected timelines for delivery.\
- Specific technical details, including metric names, label names, and label values.
Network Telemetry
The NCP shall provide network telemetry across the following domains:
- North-South (Front-End) Network (client-facing and external interconnects)\
- East-West (Back-end) Network (GPU/GPU interconnects)\
- Management Network (control plane and orchestration traffic)\
- NVSwitch Fabric (intra-node GPU switching, applicable for only GB200 and beyond clusters)\
- Host Network (NIC-level and server connectivity)
Logs
DGX Cloud will require the NCP to provide logs from various network technologies, including but not limited to:
- Fabric Manager logs for the NVLink domain (where applicable)\
- Subnet Manager logs for the NVLink domain (where applicable)\
- VPC Flow logs (all ingress/egress traffic)\
- UFM Event logs\
- General Switch Logs\
- Switch syslogs\
- Switch kernel logs\
- BMC SEL logs\
- syslogs
Storage Requirements
NCP must provide shared storage solutions (where applicable) that are manageable via standard APIs and UI, including auditing rights for NVIDIA access.
Home Directory Storage
- Quota Feature: Configurable filesystem-wide limit, default user/gid quota settings, and per uid/gid overrides.\
- Accounting: Usage accounting for uid/gids must be available when the feature is enabled.
High-Speed Storage Service Requirements
High-Speed Storage Filesystem Requirements
Data Movement Systems Requirements
The Data Movement system is used to copy data from an external data source (NVIDIA, other Cloud, etc) to the NCP data center.
DGXC-Managed Storage System Deployment
For scenarios where the storage system software will be deployed and managed by DGXC rather than the NCP, the following requirements apply. These requirements enable DGXC to operate storage systems (such as high-speed parallel filesystems, capacity object storage, or block storage) using NCP-provided infrastructure while maintaining operational control.
Host Provisioning and Lifecycle
Network Transport and Fabric Visibility
Backend Switch Fabric API
The purpose of this API is to expose sufficient information about the clusterâs network topology to enable efficient scheduling, placement, and optimization of multi-node GPU workloads. Understanding the network hierarchy between compute instances and switches, as well as both intra- and inter-node NVLink domains, is essential for minimizing communication latency and maximizing throughput. Thus, this applies to North-South, East-West, and NVLink networks (not MGMT). See the appendix for a DGXC recommended reference implementation.
Transport and Networking requirements
Non-Conflicting IP Space Allocation for the DGXC Cluster
Purpose:
Ensure DGXC GPU clusters deployed in NCP can access the NVIDIA DGXC/CorpIT network directly via routing exchange. DGXC Cluster IP address must be non-conflicting with existing NVIDIA private IP space.
Connection to NVIDIA CorpIT Network
Purpose:
Provide connection from DGXC GPU clusters within NCP to NVIDIA CorpIT for internal Command & Control and admin access.

Figure: Private Cloud interconnect + VIF + BGP for CorpIT access
Connection to DGXC Storage
Purpose:
Enable high-bandwidth, end to end MACsec-encrypted (fail-closed) access between the DGXC GPU clusters within NCP and NVIDIA DGXC on-premises object storage for large-scale data movement.

Cluster Local Internet Access
Purpose:
Provide general Internet access from DGXC GPU clusters within NCP to Internet, including NVIDIA DGXC hosted services on third-party public cloud services.

Figure: Public internet for DGXC hosted Services access
Capacity and Fleet Management
This section defines the essential metrics required for standardized monitoring and reporting of fleet health in partner engagements to support operations and contractual SLAs.
Appendix
This section contains links to reference documents and implementation guides to provide additional details if NCPs need them.
Test Legend
As a part of the AI Cloud-Ready Initiative, a validation suite is provided on GitHub, containing a set of tests that can be run to evaluate an environment against the provided requirements and specifications.
The Test Details column in the following sections contains four options:
### - Github Test reference
TBD - item on the github roadmap
add - item to be added into the roadmap
INFO - a requirement which is either not testable or currently wonât be tested from the test repo.
Implementation Guidance
Reference documents provide additional information on implementing some of the above requirements.
- Network Topology Discovery: https://github.com/NVIDIA/topograph\
- Exemplar Cloud Website: https://www.nvidia.com/en-us/data-center/ai-cloud-performance/\
- Kubernetes Security guidance : https://github.com/kubernetes/committee-security-response
Other Feature Considerations (Not Required)
- Disk Cloning: Disk cloning capability (for network-attached block devices). It should be possible to clone a disk even on a running instance.