NVIDIA Requirements for AI Clouds

View as Markdown
VersionDateDescription of Change
2.1Feb 26, 2026Initial version
2.2Apr 10, 2026Update to v2.2
2.3Jun 25, 2026Update to v2.3

Purpose and Intent

This document serves three main purposes:

  1. Setting requirements for NVIDIA Cloud Partners (NCPs) delivering GPU capacity to NVIDIA
    This is the primary requirements document from NVIDIA to any NCP providing NVIDIA GPU/AI compute and software services. These requirements cover the full stack of AI cloud infrastructure services and operations needed to run NVIDIA DGX Cloud, expanding on the NVIDIA hardware reference design.
  2. Providing a reference set of requirements for the industry
    NVIDIA is publishing this document openly so that NCPs, GPU datacenter operators, and AI practitioners can use it as a reference for the capabilities a large GPU consumer requires.
  3. Defining NVIDIA’s service delivery expectations
    NVIDIA expects services to be delivered as generally available services, not as bespoke implementations built for NVIDIA alone. NVIDIA also expects operational excellence, transparent communication, and proactive engagement from all partners.

NVIDIA will consider additional services that an NCP offers or plans to offer beyond what is described here.

Service Delivery SLAs

NCPs should be able to demonstrate their ability to meet the SLAs and operational requirements below, by category, to be considered for offtake.

Service Delivery Timelines

The NCP must demonstrate API readiness and transport establishment at least 12 weeks ahead of GPU delivery. Additionally, the NCP must provide development capacity (ancillary CPU nodes only) and high-performance storage capacity with the API integrated 8 weeks before GPU and cluster delivery. At this T-minus-8-week milestone, the Data Movement Systems Requirements must be met.

SLAs and SLOs

Managed K8s

  • Control Plane SLA target: Financially-backed 99.95%+ uptime for production.

Storage

  • Performance (QoS): Must provision the requested minimum throughput and IOPS.
  • Home Directory Storage:
    • Availability: Over 99.99% availability for unplanned incidents, exclusive of scheduled maintenance.
    • Durability: Over 99.99% for any FS less than 1 PB.
  • High-Speed Storage Service Requirements:
    • Availability (SLO): Must meet 99.99% availability in a 30-day rolling SLO exclusive of maintenance.
  • High-Speed Storage Filesystem Requirements:
    • End-to-end Availability: Over 99.5% uptime per PB.
    • Durability: Over 99.999% durability per PB annually.

Operational Requirements

  • Dedicated technical specialist or engineer available to NVIDIA.
  • Slack channel monitored by a technical specialist or engineer.
  • 24x7 support available per the partner’s standard incident severity procedures, including emergency access recovery.
  • Service-impacting incidents and planned and unplanned maintenance events are communicated to NVIDIA.
  • For planned maintenance, NVIDIA can schedule maintenance windows through APIs or console tools, avoid unexpected outages, and provide feedback.
  • NCP must remediate critical vulnerabilities in a timely manner while providing transparent disclosure of any issues.

Telemetry

NCP shall deliver all required telemetry, including metrics and logs, with a latency of no longer than 120 seconds.

Testing Compliance

The requirements described in this document can largely be validated using the tests in the AI Cloud Ready test suite. See the AI Cloud Ready documentation for details on how to run the tests. Tests are added regularly, so review the Requirements Test Matrix to find which tests validate which requirements.

Exemplar Cloud Workload Performance

NVIDIA Exemplar Cloud seeks to improve performance per TCO with hardware and software recipes, references, tools, and capabilities. Run the latest publicly available release from https://github.com/NVIDIA/dgxc-benchmarking and always select the latest release version from the GitHub repository. The release must be completed successfully on one uniform hardware cluster type. Run all workloads for a given release and share the results in the template below.

Req IDFeatureMin SizeDescription
BM01Benchmarking for Exemplar CloudRun per Scalable Unit (e.g. 512 GPU cluster)Achieve within 5% of an NVIDIA provided target performance number; this should be run on every Scalable Unit (SU) handed off.

Compute and Network Provisioning

This section outlines the requirements for provisioning compute and network resources. Compute instances can be provided as either bare-metal instances (through BMaaS) or virtual machines (through VMaaS) to support the NVIDIA DGX Cloud engagement. All operations must be controlled through a fully documented and secure API; gRPC or REST is preferred. All systems are expected to scale and perform at scale.

General, Compute, and Lifecycle Management

Req IDRequirement AreaDescription
CNP01API/CLI AccessDGXC must have API or CLI access to the NCP provisioning system for: (1) Node lifecycle management (create, update, delete, list, or manage power states (reboot, on/off power cycle); (2) Network configuration; (3) Inventory and topology discovery; (4) security configuration (users, service accounts, groups, roles); (5) Maintenance and operations (see later section)
CNP02Declarative Resource InterfacesFor resources requiring multiple steps and a workflow, please provide the appropriate mechanism. A terraform provider is preferred. E.g. automating filesystem provisioning
CNP03NVLink-Aware AllocationFor NVL72 the API must support NVLink domain-aware allocation.
CNP04Resource StatesMust support clear resource (e.g. instance, network) states where applicable. For example, provisioning, running, degraded, maintenance required, stopping, stopped, terminating, terminated.
CNP05TaggingSupport for user-defined tags/labels and cloud-init metadata on instances.
CNP06Console AccessSerial console access is required (read-only sufficient, interactive preferred).Serial console output shall be logged and be available for historic queries (at least 1 day retention, 1 month desired).
CNP07If VMaaS Present: # VMs/NodeGPU Nodes: no more than one VM per Node
General Purpose CPU Nodes: More than one per node, with ability to select via memory/core count shape.
CNP08Stable IdentifiersAll resources (e.g. nodes, switches) must have a stable and persistent ID that does not change during the lifespan, even when it goes offline for a service event. VMs must also have a stable identifier.
CNP09FirmwareBetween tenants, all firmware must be brought to a known good state, all firmware must be cryptographically signed and attested during boot.
CNP10Remote ManagementPlatform management solutions (e.g., BMC) must support Redfish over TLS (Disable IPMI).

Boot Process and Disks

Req IDRequirement AreaDescription
BOOT01Image Deployment & UpdatesAPI-driven workflow allowing DGXC to deploy, update, and manage vendor-provided or custom disk images via bare-metal, VM, or k8s node pool provisioning.
BOOT02Access to Instance Metadata from Guest OSSupport for cloud-init and instance metadata discovery via link-local addresses or virtual devices.
BOOT03Custom Disk ImagesSupport for tenant created custom OS images (either of: raw, qcow2, etc). API calls: get, list, create, delete. Images should be accessible across all tenant projects/clusters/environments.
BOOT04Node Local StorageGPU and CPU nodes support access to node local storage (NVMe / SSD) for use as scratch storage or for caching services.

SDN and Virtual Networking

This section covers the virtual networking requirements. Physical transport and networking requirements are discussed later in this document.

Req IDRequirement AreaDescription
SDN01Virtual NetworkingFull API/CLI lifecycle management (Create, Read, Update, Delete, List) for software-defined private networks. Must support non-conflicting BYOIP (including 7.0.0.0/8) and stable private IP allocations. Applicable to all types of resource nodes (CPU, GPU, Storage, etc).
SDN02Security GroupsSupport for VPC-style security groups (or equivalent), including IP/CIDR-based allow and deny rules. Must define scope/application at workload, node, service (e.g. K8s API Service) and subnet/tenant levels.
SDN03Security OperationsFull API/CLI capabilities to Create, Read, Update, and Delete security groups, including defined audit processes
SDN04Tenant IsolationHard logical or physical network segmentation for out-of-band management (BMC), user traffic, and storage-specific operations.
SDN05Floating/Movable IPAbility to automatically or API-driven switch a floating private IP between nodes via API within <10 seconds without requiring an instance reboot.
SDN06Localized DNSSupport for tenant-defined localized DNS configuration to enable internal domain resolution to private endpoints (e.g. storage endpoints)
SDN07VPC PeeringSupport for cross-virtual-network connectivity with full bandwidth and no “hairpin” routing.
SDN08Storage Mesh ConnectivityThe virtual network (from SDN01) must provide unrestricted L3 routing between all storage hosts, enabling full-mesh, all-to-all communication across different subnet (w/o going thru a gateway)
SDN09ObservabilityThe platform shall provide comprehensive logging for network infrastructure, including hardware faults, latency/performance fluctuations, and a detailed audit trail of all configuration changes to network filtering rules..
SDN10DNS Private DomainMust allow each nodes DNS resolver to forward a tenant-defined private domains (e.g. *.nvidia.com) to a tenant specified DNS server.

Kubernetes as a Service (KaaS) Requirements

Kubernetes Conformance, Versioning, and Compliance

Req IDRequirement AreaDescription
K8S01Certified VersionsCertified Upstream Versions: Official CNCF-certified versions only; no proprietary forks, and passes the standard Kubernetes conformance tests.
K8S02Version UpdatesSupport the three most recent minor releases (in the maintenance window); new minor versions must be available within 4-6 weeks of the upstream release; automated control plane security patching.
K8S03EOL PolicyDefined notification periods for version deprecation.
K8S04Kubernetes Security ResponseMust participate in the Kubernetes Security Response Committee (SRC) process. Must be attempting to join if not part of the security committee. Must be able to:
  • Responsibly disclose any discovered vulnerabilities to the Kubernetes SRC
  • Receive and respond to embargo notifications from the SRC
  • Patch disclosed vulnerabilities in the managed service during embargo prior to public disclosure and in compliance with direction provided from the Kubernetes SRC ensuring that the patching process does not violate embargo or SRC guidance.

Kubernetes Operational Excellence

Req IDRequirement AreaDescription
K8S05Lifecycle Management - Control PlaneAPI/CLI/Terraform Provider for CRUD provisioning; <30 min control plane bring-up.
K8S06Lifecycle Management - Node Pool
  • API/CLI/Terraform for CRUD provisioning ( e.g., create node pool, update node pool, delete node pool, scale a node pool to a target count).
    • Must be able to specify node type (specific CPU or GPU instance type) including CPU-only node pools with high-performance networking for data movement and ingest workloads
  • Ability to specify default node labels and node taints within a node pool when a node joins the cluster.
  • When down-scaling a node pool, ability to down-scale bad/specific nodes.
K8S07API Server MetricsShare API Server metrics in a Prometheus scrapable format to allow NVIDIA to measure API Server SLO.
K8S8VersioningProvider-managed control plane upgrade processes.
K8S9Zero-Downtime UpgradesMinor version control plane updates without app downtime or maintenance windows.
K8S10Node Upgradesuser-initiated rolling updates respecting pod disruption budgets.
K8S11HA Control PlaneRedundant architecture with etcd separation.
K8S12Backup & Disaster RecoverySupported recovery within defined RPO/RTO; needs to be auditable & testable

Robust Kubernetes Security

Req IDRequirement AreaDescription
K8S14Control Plane IsolationPer tenant k8s control plane nodes must be separate from worker nodes and outside of the tenant cluster/VPC.
K8S15Access ControlsCluster endpoint must provide network access controls.
K8S16IAM IntegrationKubernetes Service Accounts shall integrate with the platform IAM system to enable workloads to assume platform-managed identities and roles with appropriate scopes.
K8S17Service AccountsKubernetes shall support standard Service Accounts and projected tokens as the workload identity mechanism, including a cluster-specific OIDC issuer to enable workload identity federation. The cluster shall expose OIDC discovery and JWKS endpoints that are reachable by configured external identity consumers (e.g. AWS IAM, GCP workload identity)
K8S19EncryptionAt-Rest Encryption for etcd and secrets.
K8S20LoggingAbility to view or export Kubernetes control plane logs (apiserver, kcm).

Kubernetes Component and Extension Requirements

Req IDRequirement AreaDescription
K8S21API ExtensionsMandatory support for CRDs and Validating/Mutating Admission Controllers.
K8S22CNIStandard compliance; supports Network Policies; IPv4/IPv6 dual-stack desired.
K8S23CSI
  • NCP provides CSI Driver installable by NVIDIA (helm or kustomize) for Block, shared FS, and NFS.
  • Support for static and dynamic provisioning, snapshots, and resizing via PVs and PVCs.
  • CSI credentials are tenant cluster scoped (no cross cluster).
  • APIs to query storage usage against overall cluster quota with per PVC/Volume usage to manage utilization across PVCs and manage quotas using provided credentials
  • Vendor provided storage kernel modules and tools provided via (1) installed by CSI driver, (2) pre-installed in NCP provided machine image or (3) installable packages provided
K8S24DRAEnabled Dynamic Resource Allocation (DRA) regardless of upstream feature status (Beta/GA).
Some DRA features require enabling feature gates for the control plane, in case our customers want to run AI workload with new DRA features.
K8S25Operator SupportSupport standard operator-based management of hardware accelerators and associated drivers. Provider-default accelerator operators and drivers shall be replaceable or overridable to allow installation of tenant-required operator and driver versions (e.g., GPU Operator, Network Operator). Provide golden configurations for GPU Operator and Network Operator.

Kubernetes Functionality

Req IDRequirement AreaDescription
K8S26ClustersSupport multiple clusters in the same tenancy; support multiple clusters in the same VPC.
K8S27Kubernetes Control Plane Size PinningPin Control Plane instances to handle a particular load-limit.
K8S28PerformanceMeet the standard Kubernetes performance test certified up to 5000 nodes (or to the maximum size of the cluster, whichever is smaller) - size CP as necessary. Managed Kubernetes Control Plane SLO and Performance meets or better than the Kubernetes standards results.
K8S29Kubernetes LoadBalancer Service SupportThe platform shall support Kubernetes Service resources of type LoadBalancer, including:
  • External load balancers with publicly routable IPs
  • Internal load balancers with private IPs reachable via private network access
  • Static IP assignment
K8S30DNS ConfigurationThe platform shall support configuring Kubernetes internal DNS (e.g. CoreDNS) with conditional forwarding rules for specified DNS zones to designated enterprise or internal DNS resolvers.
K8S31Configurable Kubernetes CIDR RangesAbility to configure Kubernetes service IP range, Node IP range, and Pod IP range.

Security and Identity Management

Identity and Access Management (IAM)

The platform must provide a centralized system for authentication, identity federation, authorization, and lifecycle provisioning across all platform services. It must integrate with a trusted external or platform-hosted identity provider and consume OIDC-based identity tokens for user authentication. Upstream identity sources and protocols, such as enterprise directories, may be used through federation with the identity provider.

Req IDRequirement AreaDescription
SEC01AuthenticationUsers: Support standards-based user authentication via OIDC and SAML 2.0, including federation with external identity providers (e.g. NVIDIA’s enterprise IdP) for single sign-on (SSO) across platform and tenant-facing services. Validate OIDC-issued tokens including signature, issuer, audience, expiration, and required claims for identity and authorization decisions.
SEC03AuthenticationExternal Services: Support authentication of out of cluster service accounts for service-to-service access. Must support credential-based access, including long-lived credentials where required. If long-lived credentials (e.g. API keys) are issued, the platform will support configurable expiration and rotation. Need ownership attribution for all service accounts. The platform shall provide account information (such as detection of unused accounts).
SEC04Authorization (RBAC)The platform shall enforce least-privilege RBAC for all managed services and infrastructure, featuring granular API actions (e.g. CRUD), scopes (e.g. dev vs staging vs prod), and function (e.g. image builder, provisioner, auditor). Roles and permissions shall be assignable to groups, with users inheriting access through group membership (GBAC); group membership may be sourced from OIDC claims and/or SCIM-provisioned groups.
SEC05Identity / Directory ServicesThe platform shall integrate with the NVIDIA LDAP (RFC2307bis) directory service such that users identities and group membership can be resolved by dependent services for authentication and authorization decisions (e.g. storage - POSIX-based access control )
SEC06Workload/Service IdentitySupport standard workload, service, and node security identities using short-lived credentials, including OIDC-based workload identity federation and Kubernetes Service Accounts where applicable.
SEC07Admin InterfacesAll administrative interfaces—whether UI, CLI, or API—must be protected by Multi-Factor Authentication (e.g. mgmt API)
SEC08Audit LogsAudit logs must be generated and retained for all security-relevant events, including management and control plane API calls, authentication events, and authorization decisions. Audit logs shall be retained for a minimum of 30 days and accessible to authorized platform operators. Must provide a log export mechanism (such as publishing to an S3 bucket). Exported logs should include sufficient metadata to identify tenant, project/account, region, service, resource identifier, actor, event timestamp, source IP where applicable, action, and authorization result.
SEC23ProvisioningThe platform shall support SCIM 2.0 for automated user and group lifecycle management from enterprise identity providers. SCIM endpoints shall require authenticated and authorized access, support core User and Group resource operations, and synchronize group membership changes with the platform authorization engine. Synchronized groups shall be first-class RBAC objects targetable by role bindings and IAM policies; membership changes shall propagate promptly across all managed services.
SEC24AuthenticationMust support domain-based IdP routing, mapping multiple email domains to a designated identity provider (e.g., nvidia.com and nvw.nvidia.com to the NVIDIA enterprise IdP)
SEC25Organization-Level PoliciesThe platform shall support organization-level security guardrail policies that cascade across all subordinate tenant resources (networks, clusters, storage, compute) and cannot be weakened or bypassed by lower-level configuration. Policy violations shall be denied at resource creation/update time and recorded in audit logs.
SEC26SSO EnforcementThe platform shall allow authorized administrators to enforce federated SSO for a tenant, restricting local username/password and other non-federated login for regular users. Enforcement shall apply consistently across UI, CLI, API, and administrative interfaces.
SEC27Account ManagementThe platform shall expose a programmatic mechanism to create and manage:
  • Isolation Units (e.g., projects, sub-project)
  • IAM Users
  • Service Accounts
  • Logs

Cryptography and Key Management

Req IDRequirement AreaDescription
SEC09Key & Certificate LifecycleThe platform shall support secure issuance, distribution, storage, rotation, and revocation of cryptographic keys and certificates used across platform services. It shall support automated rotation of provider-managed and customer-managed keys and certificates, with configurable rotation intervals. Must be auditable. Must have an expiration date.
SEC10Key UsageThe platform shall support use of managed keys and certificates across platform services for encryption, authentication, and signing.

Network Isolation and Encryption

Req IDRequirement AreaDescription
SEC11Tenancy ModelHard physical or logical isolation for network, data, and compute. Separation of control planes and tenants is mandatory. This includes separation of storage resources. Provide hierarchical tenancy (at least organization → project).
SEC12BMC SecurityOut-of-band management (BMC) must be on a dedicated, restricted network (physically separate or VLAN/VRF-isolated). Direct access from the public internet or general corporate networks must be blocked, and only accessed via a hardened bastion (jumphost) server.
SEC13Network Traffic EncryptionEncryption and mutual authentication (mTLS or equivalent) for all east-west and north-south network traffic

Edge Network Security

Req IDRequirement AreaDescription
SEC14Private AccessNo public internet access by default; all API endpoints (e.g. K8s API Server) must be restricted via firewall/private link.
SEC15Edge Network Security PolicyAll traffic must be filtered via Security Groups and/or user customizable ACLs using 5-tuple rules.
SEC16EnforcementNCP must specify the enforcement technology (e.g., Hardware firewalls, SDN, DPUs/SmartNICs) and its specific placement in the packet path.
SEC17Threat Intelligence & ScaleAbility to subscribe to GeoIP threat & Embargo feeds and import them into security groups. NCP should share the max supported records/rules.
SEC18MACsec Protection LinksProtect links between NCP Data Center and NVIDIA POP.

Hardware Security and Compliance

Req IDRequirement AreaDescription
SEC19SOC 2SOC2 type 1 or better is required covering Security, Availability, and Confidentiality across all services and DC infrastructure.
SEC20At-Rest Data ProtectionMandatory encryption of all data at rest (e.g. local NVMe/SSD, network-attached storage) via Self-Encrypted Drives (SED).
SEC21Data SanitizationData sanitization must be performed between tenants or on a hardware replacement, including cryptographic erase of all data drives between tenants; sanitization/wipe of any persistent or volatile memory including SRAM/GPU memory; resetting of TPM and BIOS.
SEC22Root of Trust + Secure BootMandatory support across all platforms for Hardware Root of Trust mechanisms (TPM 2.0). The platform must enable UEFI OS Secure Boot w/ TPM 2.0.

Breakfix Requirements

The NCP must provide a specific Breakfix API to support fleet reliability. Any node-level remediation must not impact other parts of the tenancy. Specifically, NVLink must be reconfigured properly to take a node out of the tenancy.

The API must enable the following actions:

Req IDRequirement AreaDescription
BFX01Breakfix LifecycleCompute: Power-cycle individual nodes or reset a VM instance.
GPU: Reset GPUs on an individual node (as needed - k8s).

Maintenance: Return/Report an individual node and a rack to the Provider for maintenance.

Cordon: Mark a node as unschedulable for new workloads (but finish existing).

Replace: Request a host-replacement when health thresholds are breached
BFX02Breakfix Events
  • Query for any upcoming/current maintenance events for a node or rack
  • Query for any retirement notices for a node/rack.
  • Query for historical / status information for equipment repair.
  • Event information should include:
  • ticket open date
  • ticket update date
  • ticket close date
  • Hardware Stable Identifier (e.g., node ID)
  • Hardware category/type impacted (e.g., GPU, fan, interconnect)
  • Maintenance/Error/fault description (some short description of the issue)
  • Action: Categorization of action (e.g. repairs done on faulty GPUs to resolve the fault)
  • Provider Account ID
  • ticket ID
  • Node Handover Date (Date when the node was deployed in Production)
BFX03Diagnostics
  • Identify serial numbers of installed hardware (chassis, baseboard, network adapters, CPU, GPU, etc). Obfuscated but stable identifiers are also OK.
  • Inspect firmware versions of compute nodes and NV switch trays.

Telemetry Requirements

The telemetry requirements comprise two core components that require alignment between DGX Cloud and the NCP:

  1. Delivery Method: How telemetry will be delivered by the NCP to DGX Cloud for ingestion.
  2. Telemetry Scope: What telemetry the NCP will deliver to DGX Cloud.

Delivery Method

NCPs must deliver all required telemetry, including metrics and logs, in a manner that allows ingestion into DGX Cloud systems with a latency of no longer than 120 seconds. Native OpenTelemetry Protocol delivery is preferred.

Telemetry Scope

DGX Cloud will provide the NCP with a specification document with the required metrics and logs. Upon receipt, the NCP must provide a formal written response detailing:

  • Confirmation of its ability to deliver the specified metrics and logs.
  • Projected timelines for delivery.
  • Specific technical details, including metric names, label names, and label values.

Network Telemetry

The NCP must provide network telemetry across the following domains:

  • North-South (front-end) network, including client-facing and external interconnects.
  • East-West (back-end) network, including GPU/GPU interconnects.
  • Management network, including control-plane and orchestration traffic.
  • NVSwitch fabric, including intra-node GPU switching, applicable only for GB200 and later clusters.
  • Host network, including NIC-level and server connectivity.

Logs

DGX Cloud will require the NCP to provide logs from various network technologies, including, but not limited to:

  1. Fabric Manager logs for the NVLink domain, where applicable.
  2. Subnet Manager logs for the NVLink domain, where applicable.
  3. VPC Flow logs for all ingress and egress traffic.
  4. UFM event logs.
  5. General switch logs.
  6. Switch syslogs.
  7. Switch kernel logs.
  8. BMC SEL logs.
  9. Syslogs.
  10. Management logs.

Storage Requirements

NCPs must provide shared storage solutions, where applicable, that are manageable through standard APIs and UIs, including auditing rights for NVIDIA access.

Home Directory Storage

  • Quota Feature: Configurable filesystem-wide limit, default user/gid quota settings, and per-uid/gid overrides.
  • Accounting: Usage accounting for uid/gids must be available when the feature is enabled.
Req IDRequirement AreaDescription
DIR01File Service UID/GID Quota FeatureConfigurable filesystem-wide limit, default user/gid quota settings, and per uid/gid overrides available.
Usage accounting for uid/gids when the feature is enabled.
DIR02Must Be NFS Storage
  • NVIDIA requires NFSv4 protocol shared storage to work.
  • Access control based on DLs requires POSIX.
DIR03SnapshotsThe file system must support the ability to provide snapshot / restore functionality.
DIR04LDAPFile Service must support integration with an NVIDIA-managed LDAP (see SEC05)

High-Speed Storage Service Requirements

These are the requirements for provisioning and interacting with the provider’s service offering.

Req IDRequirement AreaDescription
HSS01Provisioning APIsStorage provisioning may be via vendor portal/API or NCP portal/API.
HSS02PerformanceMust provision needed throughput requested for minimum bandwidth and IOPS.
HSS03IntegrationK8s: CSI support
Breakfix API required to report storage issues
HSS04Quota SupportConfigurable filesystem-wide limit, default user/gid quota settings, and per uid/gid overrides available. Usage accounting for uid/gids when the feature is enabled.
Configurable directory quota settings … it must be possible to apply a quota for a given directory. Usage accounting for directory quotas when enabled.
HSS05Upgrade, MaintenanceProvider / NCP initiates desired maintenance.
NVIDIA can schedule actual maintenance and can defer maintenance up to 2 weeks.
Upgrades should be non-disruptive.
HSS06RDMA Memory ProtectionStorage systems using RDMA must enforce memory protection via authorization keys for both local and remote access.

High-Speed Storage Filesystem Requirements

These capabilities are required for the high-speed filesystem.

Req IDRequirement AreaDescription
HSS07Parallel High Speed FilesystemParallel or multi-path high-speed filesystem that supports scaling to thousands of simultaneous clients while sustaining requested performance.
HSS08Single File System SizeIt must be possible to allocate a file system of at least 1 PiB even if the initial request is less. Growing to > 10PiB as cluster size increases.
This hard requirement may be higher for a specific site and if so will be communicated via the ancillary services document.
HSS09Multiple Filesystems (Fungible Total Capacity)Can have >1 filesystem within our total capacity. Minimal file system size <= 50 TiB.
HSS10Filesystem ExpansionLive file system expansion is supported, in terms of capacity, inodes, IO performance, and metadata operations performance. Performance should scale linearly with capacity.
HSS11Client
  • Ability to describe your client: In-Kernel, userspace, or bare-metal client installation requirements.
  • Support integration with client kernels / OS used by NVIDIA, as needed.
  • DKMS-enabled packages available for Ubuntu 20.04, 22.04, and 24.04-based operating systems.
  • ARM64 versions compatible with GB200-ready kernels are mandatory, e.g. Linux 6.8.x.
    Managed Storage Service Provider will provide client configuration best practices and configuration guidelines for filesystem options and kernel module configuration to reliably achieve optimal performance on ARM and x86_64-based clients.
HSS12Quota (User, Project & Group)Must support soft and hard quotas - uid / gid / project(directory)-id quotas with enforcement.
HSS13Root-SquashNvidia needs to be able to enable or disable and manage root-squash at any time.
HSS14FlockIt must be possible to mount the file system with flock.
HSS15Ability to Audit ChangesEnable Nvidia to have access to changelog data for filesystem auditing and detailed user operations tracking.
Tracking by uid/gid, create files, create dirs, rename files, rename dirs, delete files,
delete dirs
HSS16HAAll services are required to tolerate any critical component failure in the backend and provide continued client access to all storage services in such cases.
HSS17Multi-Node CoherencyOne second or less for client attribute and dentry cache updates/invalidates
HSS18Client MultipathingClients must have multipathing to all storage servers.
HSS19LDAP (for NFS)NFS-based high-speed filesystem services must support integration with an NVIDIA-managed LDAP server (including unix uid group membership for users with > 16 group memberships) as per SEC05.

Data Movement Systems Requirements

The Data Movement system copies data from an external data source, such as NVIDIA or another cloud, to the NCP data center. These requirements must be met eight weeks before the first tranche of GPU delivery.

Req IDRequirement Description
DMS01Dedicated K8s ClusterProvider-managed k8s cluster (or ability to stand up our own) for Data Mover stack available ahead of the GPU cluster bringup to pre-stage data
DMS02Data Mover Nodes (CPU)Dedicated CPU nodes for running data mover - needs high performance networking (exact quantity will be communicated via ancillary services doc)
DMS03Access to Same GPU StorageSame filesystem as mounted on GPU nodes mounted on the Data Mover nodes (or ability to mount the same filesystem via CSI)
DMS04Access to NVIDIA Corp NetDedicate link (as described in network transport) to NVIDIA corp net, preferably with vpn, but otherwise with stable IP for allowlisting.
DMS05Stable Egress IPStable IP to IP allowlist access to Nvidia services. (e.g. similar to NAT Gateway)

DGXC-Managed Storage System Deployment

For scenarios where DGXC, rather than the NCP, deploys and manages the storage-system software, the following requirements apply. These requirements enable DGXC to operate storage systems, such as high-speed parallel filesystems, capacity object storage, or block storage, using NCP-provided infrastructure while maintaining operational control. For storage systems provided by the NCP, disregard this section.

Host Provisioning and Lifecycle

Req IDRequirement AreaDescription
STG01Operating System SupportNCP must support a workflow that allows DGXC storage operators to integrate vendor-provided or storage-specific operating system images via bare-metal or VM provisioning for storage servers. The workflow must: (a) Allow DGXC to deploy custom OS images (e.g., vendor-enhanced kernels for Lustre, Rocky Linux, Ubuntu 20.04/22.04/24.04).
STG02Drive Sanitization PolicyCryptographically erase data drive contents between storage system tenants with full attestation of host firmware. Must support an optional flag to skip drive sanitization during break/fix flows (e.g., power supply replacement) where tenancy does not change. Critical hardware component replacements may require sanitization without override, this is inclusive of GPU / CPU node local storage.
STG03Stable IP AssignmentStorage nodes must support static IP addressing that remains stable during host lifecycle operations and does not reset between maintenance events.
STG04Out-of-Band Failure DetectionNCP must provide the ability to detect system failures out-of-band, including device, network, memory, and drive failures, enabling DGXC to proactively respond to hardware issues.
STG05Topology ObservabilityNCP must provide visibility into failure domains to enable DGXC to provision storage nodes with physical diversity. Storage systems must be able to provision nodes that purposefully span failure domains for resilience.
STG06BlueField/DPU SupportFor storage systems utilizing BlueField-based architectures, the host provisioning system must support lifecycle management and specific configuration requirements for BlueField “JBOF” systems that export NVMe-oF to hosts.

Network Transport and Fabric Visibility

Backend Switch Fabric API

The purpose of this API is to expose sufficient information about the cluster network topology to enable efficient scheduling, placement, and optimization of multi-node GPU workloads. Understanding the network hierarchy between compute instances and switches, as well as intra- and inter-node NVLink domains, is essential for minimizing communication latency and maximizing throughput. This applies to north-south, east-west, and NVLink networks, but not to management networks. See the appendix for a DGXC-recommended reference implementation.

Req IDRequirement AreaDescription
NET01Backend Switch Fabric APIFor each compute node, the API must provide visibility into the backend network switches connecting the node to the core.
  • Identification: Each switch must be identified by a unique, stable identifier. A “switch” may represent a physical switch or a logical connectivity domain.
  • Structure: API may be gRPC or REST. Response structure may include multiple nodes (pagination expected).
  • Topology: Switch info can be returned as an ordered array of IDs (e.g., leaf, spine, core) or separate fields for each tier.
NET02NVLink Domain APIRequirement: For compute nodes supporting NVLink (e.g., GB200, GB300, Vera Rubin), the API shall return the unique identifier of the NVLink domain associated with each node.

Implementation: Can be a separate API method or part of the Backend Switch Fabric API.

Transport and Networking Requirements

Non-Conflicting IP Space Allocation for the DGXC Cluster

Purpose: Ensure DGXC GPU clusters deployed in an NCP can access the NVIDIA DGXC/CorpIT network directly through routing exchange. DGXC cluster IP address space must not conflict with existing NVIDIA private IP space.

Req IDRequirement AreaDescription
NET03Non-Conflicting IP Space Allocation for the DGXC ClusterBring Your Own IP (BYOIP): NCP shall support the ability for NVIDIA to bring and allocate its own IP private address space for DGXC GPU clusters.

Stable IP: NCP shall provide a possibility to create static IP allocations that persist across instance restarts and re-creations. That includes floating IP allocations.

DoD space: NCP shall support allocation and use of the 7.0.0.0/8 IPv4 address space for DGXC GPU cluster deployments. This IP space shall be considered equivalent to RFC1918 addresses

Routing Support: NCP must support advertising and routing of BYOIP prefixes within the NCP environment and across interconnects (Private Cloud Interconnect, IPSec, etc.)

Connection to NVIDIA CorpIT Network

Purpose: Provide a connection from DGXC GPU clusters within the NCP to NVIDIA CorpIT for internal command, control, and administrative access.

Req IDRequirement AreaDescription
NET04Connection to NVIDIA CorpIT NetworkBandwidth: Low bandwidth (Up to 10Gbps).

Transport: Private Cloud interconnect + VIF + BGP (preferred for better performance/security). DGXC will establish connectivity to NCP through a mutually agreed Point of Presence (POP) using Private Cloud Interconnect, functionally equivalent to AWS Direct Connect, GCP Dedicated Interconnect, Azure ExpressRoute, and OCI FastConnect. Connectivity will be provisioned with a Virtual Interface (VIF) and routing established via BGP. The interconnect will be used to exchange private IP space (RFC1918, as well as 7.0.0.0/8) between DGXC and NCP.

Corporate network connectivity diagram

Figure: Private Cloud Interconnect + VIF + BGP for CorpIT access

Connection to DGXC Storage

Purpose: Enable high-bandwidth, end-to-end MACsec-encrypted, fail-closed access between DGXC GPU clusters within the NCP and NVIDIA DGXC on-premises object storage for large-scale data movement.

Req IDRequirement AreaDescription
NET05Connection to DGXC StorageTransport: Private Cloud interconnect + VIF + BGP (preferred for better performance/security). DGXC will establish connectivity to NCP through a mutually agreed Point of Presence (POP) using Private Cloud Interconnect, functionally equivalent to AWS Direct Connect, GCP Dedicated Interconnect, Azure ExpressRoute, and OCI FastConnect. Connectivity will be provisioned with a Virtual Interface (VIF) and routing established via BGP. The interconnect will be used to exchange private IP space (RFC1918, as well as 7.0.0.0/8) between DGXC and NCP.

Storage connectivity diagram

Cluster Local Internet Access

Purpose: Provide general internet access from DGXC GPU clusters within the NCP to the internet, including NVIDIA DGXC hosted services on third-party public-cloud services.

Req IDRequirement AreaDescription
NET06Cluster Local Internet AccessCluster Internet access: Egress NAT IPs should be a static pool dedicated to only Nvidia Cluster/Tenancy/VPC. These persistent IP addresses must be used exclusively for DGXC traffic and shall not be shared with or carry traffic from other NCP tenants.

Availability: Must support redundant upstream paths to ensure connectivity under failure.

Internet access diagram

Figure: Public internet access for DGXC-hosted services

Capacity and Fleet Management

This section defines the essential metrics required for standardized monitoring and reporting of fleet health in partner engagements, supporting operations and contractual SLAs.

Req IDRequirement AreaDescription
CAP01Governance MetricsRequired Governance Metrics
The core metrics needed to track fleet health are:
  • Delivered: Nodes/GPUs provisioned and available to NVIDIA, allocated to a specific account/project/tenant.
  • Healthy: Nodes/GPUs functioning and meeting SLA requirements, allocated to a specific account/project/tenant.
  • Reserved: Resources allocated to a specific account/project/tenant.
  • Total Active/In-Use: Nodes/GPUs currently in use within a specific account/project/tenant.
CAP02Resource Governance API MetricsThe Resource Governance API must return the following information for each node:
  • Node ID (Unique identifier for a GPU node)
  • Health State (Healthy/Unhealthy classification)
  • Instance ID (Identifier for virtual workload)
  • Creation Timestamp (Time workload/node was created)
  • Hardware Type (Descriptor for the hardware model)
  • GPU Count (Number of GPUs per node)
  • Top-levelAccount/ID (Identifier for the top-level organization/account)
  • Sub-LevelProject/ID (Identifier for the nested project/sub-account)
  • In Use (True/False status indicating if the GPU Node is turned on and in use)
  • Region (Region of the data center where nodes are deployed)
CAP03Resource Discovery APIsIt is not acceptable to have capacity be “handed” to DGXC through a phone, slack or email message. For example, when cluster first comes online, nodes/racks are likely being handed off weekly (or more frequently). Instead, please provide the following mechanism (and we can poll):

Programmatic Capacity Discovery: All newly delivered capacity must be discoverable via a centralized API. This “Resource Index” must provide a stable resource identifier and some information on why it’s being provided (e.g. capacity fulfillment on gb300 project, break-fix / RMA return to cluster, etc)
CAP04Logical Compartmentalization & Resource IsolationTo ensure performance consistency and security, the NCP must support strict logical and physical isolation of NVIDIA’s reserved capacity.

Capacity Reservations: A mechanism to logically group and “pin” a set of resources (compute, network, storage) to accounts (or equivalent constructs) in an NVIDIA tenancy

Atomic Allocation: Support for reserving a “topology block” as a single unit, ensuring all resources in that block share identical performance characteristics and security boundaries.
CAP05Unified Health & Lifecycle APIsNVIDIA requires a “single source of truth” for the health of both physical hosts and logical compute primitives.

Per-Host Health: Real-time API access to the health bits of physical hardware (GPU state, thermal status, memory health).

Primitive-Level Status: Health aggregation at the cluster, nodegroup, or reservation level to identify broad infrastructure failures (e.g., a spine switch failure affecting a whole block).

Appendix

This section contains links to reference documents and implementation guidance that provide additional details for NCPs.

Implementation Guidance

The following reference documents provide additional information on implementing some of the requirements in this guide.

  1. Network Topology Discovery: NVIDIA/topograph. Aligning with this YAML format may be useful. Topograph currently provides in-cluster topology.
  2. Breakfix B200: B200 DGXC Lazarus BreakFix Requirements.
  3. Breakfix GB300: GB300 DGXC Lazarus BreakFix Requirements.
  4. Breakfix scenarios: DGXC Breakfix Maintenance Events.
  5. Networking: Revised GNI - NCP - Cluster Connection Requirements for DGX Cloud.
  6. Exemplar Cloud: NVIDIA Exemplar Cloud and the DGXC benchmarking repository.
  7. Kubernetes security guidance: Kubernetes Security Response Committee.

Other Feature Considerations (Not Minimum Requirements)

  1. Disk Cloning: Disk-cloning capability for network-attached block devices. It should be possible to clone a disk even on a running instance.
  2. Managed Control Plane Autoscaling: Strong preference for the control plane to automatically add capacity when load increases.
  3. Threat Detection: Control planes, management planes, and hosts under the service provider’s control should deploy threat- and anomaly-detection solutions, for example HIDS and NIDS, that can identify active threats.
  4. Break-Glass Administrative Access: The platform should support a limited break-glass access mechanism for designated NVIDIA administrative users when federated SSO is unavailable, misconfigured, compromised, or otherwise prevents authorized access to the tenant. Break-glass accounts should use local platform credentials independent of the external identity provider, be explicitly excluded from SSO enforcement, and be protected by strong authentication controls including MFA. Their use should be auditable, generate security-relevant logs and alerts, and support periodic review, rotation, disablement, and testing.