> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/dsx/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/dsx/_mcp/server.

# Inference Provider Platform Requirements on GB300 NVL72 for NCPs

> Requirements for an NVIDIA Cloud Partner (NCP) building a GB300 NVL72 compute cluster that minimizes friction for Inference Service Providers (ISPs) deploying and running inference platforms.

## Purpose

This document specifies the requirements for an NVIDIA Cloud Partner (NCP) to build a GB300 NVL72 compute cluster that minimizes friction for Inference Service Providers (ISPs) deploying and running their inference platforms. The result is managed Kubernetes that the NCP delivers to the ISP, which the ISP plans to deploy its stack on.

The document assumes that cluster hardware and facility are built, and that the physical build follows the NVIDIA DSX reference designs. It defines what the NCP does from that point forward: provision bare metal, stand up networking and multi-tenancy, storage, and the managed Kubernetes platform, and expose the telemetry, break-fix, and capacity that operations depend on.

The [NVIDIA Requirements for AI Clouds](https://docs.nvidia.com/dsx/ncp/nvidia-requirements-for-ai-clouds/introduction) document is the general requirements bar for any NCP delivering GPU capacity. Broader multi-tenant build guidance is in the [NCP Software Reference Guide](https://docs.nvidia.com/dsx/ncp/software-reference-guide/introduction).

Each section has two parts: prose that provides context and a requirements table that outlines the obligations. The tables are the normative part of this document. The prose explains why a requirement exists and how the pieces fit together; it does not itself impose requirements, so it does not repeat the requirement levels. Requirement levels follow [IETF RFC 2119](https://datatracker.ietf.org/doc/html/rfc2119):

* **MUST** and **MUST NOT** are mandatory.
* **SHOULD** and **SHOULD NOT** are strong recommendations that a partner deviates from only with good reason.
* **MAY** is optional.
* Rows marked **INFO** are informational and place no obligation on the NCP.

Each table row also shows who owns the work, under the NCP-to-ISP responsibility split:

* **Owns**: the accountable party.
* **Provides**: provides the substrate capability.
* **Exposes**: exposes an interface or topology for the other party to use.
* **Consumes**: uses what the other party provides.
* **Executes**: acts on request.
* **n/a**: not applicable to that party.

Throughout, the NCP is the Operator and owns the substrate; the ISP is the Tenant and owns its inference software stack on the managed cluster.

## Reference Designs

The NCP adheres to the following reference designs as a starting point: the [DSX Facilities Infrastructure Design Guide](https://partners.nvidia.com/DocumentDetails?DocID=1152370) for facility, power, and cooling; the [GB300 NVL72 Reference Design](https://partners.nvidia.com/DocumentDetails?DocID=1138480) for cluster architecture, server specification, and network cabling; and the [storage partner requirements](https://partners.nvidia.com/DocumentDetails?DocID=1124019). These reference designs also specify the network-appliance requirements for edge connectivity.

Access to the linked NVOnline documents must be granted by an NVIDIA representative and may not be available to ISPs.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                 | Keyword  | NCP (Operator) | ISP (Tenant) |
| :----- | :---------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- | :------------- | :----------- |
| RD-1   | Adhere to the Facilities Infrastructure Design Guide (NVOnline [#1152370](https://partners.nvidia.com/DocumentDetails?DocID=1152370)).                      | **MUST** | **Owns**       | **n/a**      |
| RD-2   | Adhere to the GB300 NVL72 Reference Design (NVOnline [#1138480](https://partners.nvidia.com/DocumentDetails?DocID=1138480)) — server spec, network cabling. | **MUST** | **Owns**       | **n/a**      |
| RD-3   | Adhere to the storage partner requirements (NVOnline [#1124019](https://partners.nvidia.com/DocumentDetails?DocID=1124019)).                                | **MUST** | **Owns**       | **n/a**      |

## GPU Compute Hardware, Cluster and Capacity

The GPU compute servers must be in the Qualified Systems Catalog. On that hardware, the operating system must support Kubernetes as the compute delivery mechanism, and a supported Linux distribution is recommended; refer to the [NVIDIA AI Enterprise support matrix](https://docs.nvidia.com/ai-enterprise/release-8/latest/support/support-matrix.html) for distributions that include the driver components. This revision targets bare-metal environments, so KVM-based Kubernetes delivery using virtual machines is out of scope.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                                                          | Keyword    | NCP (Operator) | ISP (Tenant) |
| :----- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :----------- |
| CMP-1  | GPU compute servers meet [NVIDIA Certified Systems Requirements](https://www.nvidia.com/en-us/data-center/products/certified-systems/)                                                                                                                                                               | **MUST**   | **Owns**       | **n/a**      |
| CMP-2  | The OS supports Kubernetes as the compute-delivery mechanism.                                                                                                                                                                                                                                        | **MUST**   | **Owns**       | **n/a**      |
| CMP-3  | Use a supported Linux distribution (see the [NVIDIA AI Enterprise](https://docs.nvidia.com/ai-enterprise/release-8/latest/support/support-matrix-8/8.1.html#bare-metal-deployments) bare-metal deployment matrix). Ubuntu 22.04 and 24.04 are currently supported for GB300 NVL72 1.0.5/1.0.6/2.0rc. | **SHOULD** | **Owns**       | **n/a**      |
| CMP-4  | KVM-based VM delivery of Kubernetes is out of scope for this revision.                                                                                                                                                                                                                               | **INFO**   | Info           | Info         |

A cluster is defined as a collection of compute nodes, control nodes, networking, and storage mapped to a specific build captured in a single hardware bill of materials. In this document, we assume one managed cluster per region per tenant (or per requested tenant cluster).

A cluster is a collection of compute nodes, control nodes, networking, and storage mapped to a specific build captured in a single hardware bill of materials. This document assumes one managed cluster per region per tenant, or per requested tenant cluster.

### Hardware Validation

Hardware validation confirms that each node, and the cluster as a whole, operate correctly before the cluster performs production work. It applies the acceptance-testing bar from the [NVIDIA Requirements for AI Clouds](https://docs.nvidia.com/dsx/ncp/nvidia-requirements-for-ai-clouds/introduction) to GB300 NVL72, using the [NVIDIA AI Cloud-Ready Validation Suite](https://github.com/NVIDIA/ISV-NCP-Validation-Suite) for functional validation and [NVIDIA Exemplar Cloud](https://github.com/NVIDIA/dgxc-benchmarking) for performance baselines.

### NCP-Managed Inventory Services

The NCP maintains a single source of truth for the fleet: an authoritative, real-time record of every GPU, node, region, and NCP, and the current state of each. It tracks identity and location (NCP, region, rack, NVL72 chassis, node, GPU SKU), topology (NVLink Domain membership, NVSwitch fabric ID), software versions (driver, [CUDA](https://developer.nvidia.com/cuda-toolkit), NCCL, and fabric-manager state), health (DCGM), and allocation ([MIG](https://www.nvidia.com/en-us/technologies/multi-instance-gpu/) geometry, assigned pool, current workload). An NCP builds this on its own asset system or from DSX OS components that already emit these fields: [NVIDIA Infra Controller](https://docs.nvidia.com/infra-controller/documentation/home) for discovery, provisioning, and node state; [Fleet Intelligence](https://docs.nvidia.com/fleet-intelligence/latest/index.html) for health and visibility; [DCGM](https://github.com/NVIDIA/dcgm-exporter) for GPU health; and the fabric manager with topology labeling for the NVLink and NVSwitch fields.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                                                                                                                      | Keyword    | NCP (Operator) | ISP (Tenant)            |
| :----- | :--------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :---------------------- |
| HW-1   | NCCL throughput benchmarks using [nccl-tests](https://github.com/NVIDIA/nccl-tests) reach SKU performance baselines; SHARP-accelerated collectives exceed standard multi-node baselines. NVIDIA also provides a [profiler/inspector plugin](https://github.com/NVIDIA/nccl/blob/master/plugins/profiler/inspector/README.md) for NCCL that provides JSON output. | **SHOULD** | **Owns**       | **Consumes** (evidence) |
| HW-2   | IB/RoCE rail tests using [ib\_write\_bw](https://github.com/linux-rdma/perftest) on every rail, with loss and error counters clean; identify bandwidth and latency outliers across IB links.                                                                                                                                                                     | **SHOULD** | **Owns**       | **Consumes** (evidence) |
| HW-3   | [GPUDirect RDMA](https://docs.nvidia.com/cuda/gpudirect-rdma/) and [GPUDirect Storage](https://docs.nvidia.com/gpudirect-storage/) smoke tests, covering low-latency network storage in addition to NVMe read and write.                                                                                                                                         | **SHOULD** | **Owns**       | **Consumes** (evidence) |
| HW-4   | NVLink and NVSwitch link-count and bandwidth checks using [nvbandwidth](https://github.com/NVIDIA/nvbandwidth).                                                                                                                                                                                                                                                  | **SHOULD** | **Owns**       | **Consumes** (evidence) |
| HW-5   | 24-hour soak under representative inference load at a sustained power cap.                                                                                                                                                                                                                                                                                       | **SHOULD** | **Owns**       | **Consumes** (evidence) |
| HW-6   | NCCL auto-discovers topology and delivers it through a standard API in XML format.                                                                                                                                                                                                                                                                               | **SHOULD** | **Owns**       | **Consumes**            |
| HW-7   | NCP-provided Inventory Service is the single source of truth for every GPU/node/region/NCP and tracks the listed fields.                                                                                                                                                                                                                                         | **MUST**   | **Owns**       | **Consumes**            |
| HW-8   | Nodes undergo hardware attestation and acceptance tests before joining a production cluster (NICo alert classifications).                                                                                                                                                                                                                                        | **SHOULD** | **Owns**       | **Consumes** (evidence) |

## Compute Operating System

The NCP provides an infrastructure-as-a-service capability to automate Linux OS installation on compute nodes. The bare-metal lifecycle is documented in the [NICo open source project](https://github.com/NVIDIA/infra-controller/blob/main/docs/overview/lifecycle.md), and the Linux distribution support matrix is covered in the [NVIDIA AI Enterprise documentation](https://docs.nvidia.com/ai-enterprise/index.html).

**Requirements Summary**

| Req ID | Requirement                                                                    | Keyword  | NCP (Operator) | ISP (Tenant) |
| :----- | :----------------------------------------------------------------------------- | :------- | :------------- | :----------- |
| OS-1   | Provide an IaaS capability to automate Linux OS installation on compute nodes. | **MUST** | **Owns**       | **n/a**      |

## Networking and Multi-tenancy

The cluster network is built on the [GB300 NVL72 Reference Design](https://partners.nvidia.com/DocumentDetails?DocID=1138480), published to NCPs through the NVIDIA NVOnline partner portal, which also specifies the network-appliance requirements for edge connectivity. On top of that hardware, two planes carry the work: a north-south plane on the [BlueField-3 DPU](https://www.nvidia.com/en-us/networking/products/data-processing-unit/) that enforces zero trust and hosts multi-tenancy and carries the control plane, storage, and external connectivity, and an east-west plane on the back-end RDMA fabric, InfiniBand or [Spectrum-X](https://www.nvidia.com/en-us/networking/spectrumx/), that moves tensor-parallel collectives over [NCCL](https://developer.nvidia.com/nccl), KV-cache over [NIXL](https://github.com/ai-dynamo/nixl), and model weights between nodes. The NCP handles IP addressing and inter-host routing, with BGP as the routing protocol. How Kubernetes services meet the edge, and how edge connectivity is sized to inbound request volume, are covered in Kubernetes Services for Edge and the Inference Routing Data Plane.

Multi-tenancy is enforced at the network level, not at the application level. Host networking on the BlueField DPUs and the network appliances, the Spectrum Ethernet switches and [Quantum InfiniBand](https://www.nvidia.com/en-us/networking/quantum2/) switches, keep tenants isolated end-to-end, and the [NVIDIA Infrastructure Controller](https://docs.nvidia.com/infra-controller/documentation/home) pairs with the network to enforce that isolation. The design uses a BGP underlay with EVPN, VXLAN overlay encapsulation, access control lists on the DPU, and zero trust between the DPU and the compute host. Kubernetes interaction with multi-tenancy is covered in the Tenancy Model section.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                            | Keyword    | NCP (Operator) | ISP (Tenant)             |
| :----- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :----------------------- |
| NET-1  | North-south through the BlueField-3 DPU enforcing zero trust and host multi-tenancy; carries control plane, storage, external connectivity ([DOCA](https://developer.nvidia.com/networking/doca)/DPF). | **MUST**   | **Owns**       | **Consumes**             |
| NET-2  | East-west back-end RDMA fabric (InfiniBand or Spectrum-X) for multi-node serving: NCCL collectives (SHARP where supported).                                                                            | **MUST**   | **Provides**   | **Consumes**             |
| NET-3  | East-west back-end RDMA fabric that enables KV-cache transfer using NIXL, weight P2P.                                                                                                                  | **SHOULD** | **Provides**   | **Consumes**             |
| NET-4  | Dual-stack IPv4 + IPv6 on egress.                                                                                                                                                                      | **SHOULD** | **Owns**       | **Consumes**             |
| NET-5  | BGP is the routing protocol with IPv4 and EVPN/VPNv4 address families.                                                                                                                                 | **MUST**   | **Owns**       | **n/a**                  |
| NET-6  | Multi-tenancy uses BGP underlay + EVPN, VXLAN overlay, DPU ACLs, and zero trust between the DPU and the compute host.                                                                                  | **SHOULD** | **Owns**       | **Consumes** (isolation) |
| NET-7  | Networking device names use standard naming convention starting with mlx5\_\*.                                                                                                                         | **SHOULD** | **Owns**       | **Consumes**             |
| NET-8  | Label InfiniBand and RoCE accurately, and ensure consistent labeling is passed to NCCL.                                                                                                                | **MUST**   | **Owns**       | **Consumes**             |

Kubernetes interaction with multitenancy is addressed in the [Kubernetes Tenancy](#tenancy-model) section.

## Storage Architecture

Storage for a GB300 NVL72 cluster follows the [NCP Reference Design for Storage](https://partners.nvidia.com/DocumentDetails?DocID=1124019), and applies the [AI Cloud Requirements storage bar](https://docs.nvidia.com/dsx/ncp/nvidia-requirements-for-ai-clouds/storage-requirements) to GB300 NVL72. The goal is high-throughput, low-latency storage that GPU compute nodes can read from and write to directly, delivered in a way that the ISP does not manage the underlying layout.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                   | Keyword    | NCP (Operator) | ISP (Tenant)                  |
| :----- | :-------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :---------------------------- |
| STO-1  | Adhere to the NCP Reference Design for Storage ([NVOnline #1124019](https://partners.nvidia.com/DocumentDetails?DocID=1124019)).              | **MUST**   | **Owns**       | **n/a**                       |
| STO-2  | Deploy storage with RDMA.                                                                                                                     | **MUST**   | **Owns**       | **Consumes**                  |
| STO-3  | Certified Storage Platforms from Partners ([Link](https://docs.nvidia.com/certification-programs/certified-storage/latest/systems-list.html)) | **MUST**   | **Owns**       | **Consumes**                  |
| STO-4  | Validate storage performance from the GPU compute nodes before production.                                                                    | **SHOULD** | **Owns**       | **Consumes** (evidence)       |
| STO-5  | Support and implement GPUDirect Storage.                                                                                                      | **SHOULD** | **Provides**   | **Consumes**                  |
| STO-6  | Locate the local disk path at /mnt/stateful\_partition; handle SSD RAID0 natively; deliver a single pre-striped NVMe mount.                   | **SHOULD** | **Owns**       | **Consumes** (burden removed) |
| STO-7  | Customer data isolation: per-tenant prefixes with IAM-equivalent policies; per-tenant KMS encryption; BYOK option.                            | **MUST**   | **Owns**       | **Consumes** (per-tenant)     |

## Kubernetes Platform

The NCP provides managed Kubernetes. The ISP does not run its own control plane, [etcd](https://etcd.io/), or upgrade pipeline; those are the NCP's responsibility under the SLA. There is one managed Kubernetes cluster per region per ISP tenant, and this document targets Kubernetes on bare metal. This section applies the [AI Cloud Requirements KaaS](https://docs.nvidia.com/dsx/ncp/nvidia-requirements-for-ai-clouds/kubernetes-as-a-service-kaa-s-requirements) section to GB300 NVL72.

**Requirements Summary**

| Req ID | Requirement                                                                                                                              | Keyword    | NCP (Operator) | ISP (Tenant) |
| :----- | :--------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :----------- |
| KUB-1  | NCP fully manages the control plane, etcd, and upgrade pipeline under the SLA; one managed Kubernetes cluster per region per ISP tenant. | **MUST**   | **Owns**       | **Consumes** |
| KUB-2  | Version-skew policy no faster than the validated combination matrix; minimum 12-month support per minor.                                 | **SHOULD** | **Owns**       | **Consumes** |
| KUB-3  | Per-region pre-upgrade canary clusters validate driver × CUDA × NCCL × kubelet before rollout.                                           | **SHOULD** | **Owns**       | **Consumes** |
| KUB-4  | Programmatic, no-downtime node-pool replacement (surge upgrades).                                                                        | **MUST**   | **Owns**       | **Consumes** |
| KUB-5  | Kubernetes DRA support for dynamic GPU resource allocation.                                                                              | **SHOULD** | **Provides**   | **Consumes** |
| KUB-6  | Per-region quota and limit visibility through an API.                                                                                    | **MUST**   | **Provides**   | **Consumes** |
| KUB-7  | No proprietary lock-in of CNI, ingress, CSI, or scheduler; upstream/OSS usable.                                                          | **MUST**   | **Owns**       | **Consumes** |

Everything above the managed Kubernetes boundary is owned by the ISP and is out of scope for this document, which specifies only the NCP substrate. The ISP owns its workload-tier CRDs, controllers, operators, admission policies, and schedulers; any replacement or extension of the defaults the managed service ships (for example, swapping the CNI to Cilium where allowed, or adding a custom scheduler extender alongside the default); the model-serving, autoscaling, routing, and observability stack; and GitOps reconciliation against the cluster (for example, Argo CD). This document references these areas only to define the substrate capability the NCP must expose.

### Installing Kubernetes

Operational management of Kubernetes is the NCP's responsibility: selecting the distribution and versions, installing Kubernetes, and verifying the installation.

**Requirements Summary**

| Req ID | Requirement                                                                                   | Keyword  | NCP (Operator) | ISP (Tenant) |
| :----- | :-------------------------------------------------------------------------------------------- | :------- | :------------- | :----------- |
| KUB-8  | NCP selects the distribution and version, installs Kubernetes, and verifies the installation. | **MUST** | **Owns**       | **Consumes** |

### Preparing Kubernetes for Inference

NCPs and tenants can validate the components using the [AI Cluster Runtime](https://docs.nvidia.com/aicr/overview/introduction/).

AI Cluster Runtime will create an inference-focused recipe on Ubuntu targeting GB300NVL72:

```shell
aicr recipe --accelerator gb300 --os Ubuntu --intent Inference --platform Dynamo -o recipe.yaml
```

The recipe can be installed using Helm charts by creating a Helm bundle:

```shell
aicr bundle --recipe recipe.yaml --deployer helm --output ./bundles
```

This bundle includes all the necessary components for Kubernetes to run a performant Inference stack.

**Requirements Summary**

| Req ID | Requirement                                                                       | Keyword  | NCP (Operator)                | ISP (Tenant) |
| :----- | :-------------------------------------------------------------------------------- | :------- | :---------------------------- | :----------- |
| KUB-9  | NVIDIA provides an AI Cluster Runtime recipe for validating inference components. | **INFO** | Owns (Applying NVIDIA recipe) | Consumes     |

### GPU Driver Management

The [GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html) keeps the GPU software stack consistent across the fleet by managing drivers in the container runtime through the embedded [Container Toolkit](https://github.com/NVIDIA/nvidia-container-toolkit), and collecting health and telemetry from [DCGM](https://github.com/NVIDIA/dcgm-exporter). The requirements describe how it behaves on an NCP cluster: it supports GPUDirect RDMA, it allows driver upgrades without competing with other services, such as gpud, for system locks, and its driver management can be disabled when the NCP provisions drivers another way.

For GPUDirect RDMA, the NCP delivers working, validated GPUDirect RDMA; whether it uses dma-buf or nvidia\_peermem is an implementation choice set by the kernel and RDMA stack.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                                                                 | Keyword  | NCP (Operator)      | ISP (Tenant)  |
| :----- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- | :------------------ | :------------ |
| KUB-10 | [GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html) builds drivers against the current running kernel using the [Container Toolkit](https://github.com/NVIDIA/nvidia-container-toolkit) and provides [DCGM-based monitoring](https://developer.nvidia.com/dcgm). | **MUST** | Owns (GPU Operator) | Pins versions |
| KUB-11 | GPU Operator interoperates with RDMA, can be disabled, and allows driver upgrades without contending with other services (for example, gpud) for system locks.                                                                                                                                              | **MUST** | Owns (GPU Operator) | Pins versions |

### Kubernetes Networking

The NCP selects the container network interface (CNI) at install time. On a GB300 NVL72 cluster, that choice has to satisfy three demands at once: pods that scale dynamically, tenant isolation that works with the BlueField DPU and the north-south Ethernet, and access to the back-end RDMA network for GPU-to-GPU communication. The specific product is an NCP decision. The NCP can also use Network Operator to manage the GPU-to-GPU communication.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                 | Keyword    | NCP (Operator)          | ISP (Tenant)                      |
| :----- | :------------------------------------------------------------------------------------------------------------------------------------------ | :--------- | :---------------------- | :-------------------------------- |
| KUB-12 | CNI (selected at install by the NCP) supports dynamic pod scaling, tenancy interop through DPU + N/S Ethernet, and RDMA GPU-to-GPU interop. | **MUST**   | Owns (selects)          | Consumes (may swap where allowed) |
| KUB-13 | Network Operator interops with RDMA, can be disabled, and allows driver upgrades without contending with other services for system locks.   | **SHOULD** | Owns (Network Operator) | Consumes                          |

### Tenancy Model

The tenancy model establishes the boundary between the NCP, as Operator, and the ISP, as Tenant. The baseline is one NCP-managed Kubernetes cluster per ISP tenant per region, each with its own control plane, aligned to the KaaS model in the [NVIDIA Software Reference Guide](https://docs.nvidia.com/dsx/ncp/part-1-software-reference-guide/container-as-a-service). Tenants do not share a control plane for production inference, and dedicated worker hosts back each tenant, because isolation is enforced at the networking layer. That design is what supports reserved capacity, confidential computing, strict data residency, and enterprise SLOs. The managed service also exposes enough topology, quota, health, and lifecycle information for the ISP to run scheduling, autoscaling, admission control, and break-fix without owning the control plane.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                         | Keyword        | NCP (Operator) | ISP (Tenant) |
| :----- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------------- | :------------- | :----------- |
| KUB-14 | Implement the required baseline: one NCP-managed cluster per ISP tenant per region, with a dedicated control plane ([SRG KaaS model](https://docs.nvidia.com/dsx/ncp/part-1-software-reference-guide/container-as-a-service)).      | **MUST**       | Owns           | Consumes     |
| KUB-15 | Do not require multiple ISP tenants to share a control plane for production inference; instead, dedicate worker hosts to each tenant for isolation, reserved capacity, confidential computing, data residency, and enterprise SLOs. | **SHOULD NOT** | Owns           | Consumes     |
| KUB-16 | Expose enough topology, quota, health, and lifecycle information for the ISP to operate scheduling, autoscaling, admission, and break-fix without owning the control plane.                                                         | **SHOULD**     | Provides       | Consumes     |

### Kubernetes Services for Edge

Inbound inference requests cross from the edge-routed network into the Kubernetes network and land on a service running on a cluster node. To keep routing efficient, these services terminate as close to the public edge as possible. The NCP uses the Kubernetes load balancer as the standard, implemented through a custom controller that calls the load balancer's API (for example, F5, NGINX, or a custom hardware API) to create the virtual server and target group on service creation. Edge-tier services also need anycast IPs for multi-zone deployments, TLS 1.3 termination with HTTP/2 and HTTP/3, WAF, DDoS, and bot mitigation at L3 and L4, and authentication through a bearer key, JWT, or optional mTLS.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                      | Keyword    | NCP (Operator)       | ISP (Tenant)    |
| :----- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------------- | :-------------- |
| KUB-17 | Inbound edge requests terminate on a service as close to the public edge as possible.                                                                                            | **SHOULD** | Provides (edge)      | Deploys service |
| KUB-18 | Use the standard Kubernetes load balancer; implement an LB controller that calls the NCP load-balancer API (F5, NGINX, or custom) to create the virtual server and target group. | **SHOULD** | Owns (LB controller) | Consumes        |
| KUB-19 | Edge-tier services: Anycast IPs for multi-zone; TLS 1.3 termination, HTTP/2 + HTTP/3; WAF/DDoS/bot mitigation at L3/4; auth using bearer key, JWT, optional mTLS.                | **SHOULD** | Provides (edge tier) | Consumes        |

### Image Registry

Serving images are large, and pulling them across regions adds latency to every cold start and scale-out. To avoid that, the NCP provides a native in-region image registry that the ISP uses to host its images close to the compute.

**Requirements Summary**

| Req ID | Requirement                                                           | Keyword    | NCP (Operator) | ISP (Tenant)       |
| :----- | :-------------------------------------------------------------------- | :--------- | :------------- | :----------------- |
| KUB-20 | Provide a native in-region image registry for the ISP to host images. | **SHOULD** | Provides       | Consumes (can use) |

### Exposing Kubernetes Control Plane - Kubeconfig Acquisition

Once Kubernetes is installed, the ISP needs a way to reach the cluster. The NCP provides access to the Kubernetes API and [kubectl](https://kubernetes.io/docs/reference/kubectl/), gated by an IAM policy that requires authentication and authorization, so that a tenant can access only its own cluster. To make that access easier to consume, the NCP also exposes interaction with the deployed environment through an AI Server. Today, each NCP delivers kubeconfig in its own way, which is a documented source of friction for ISPs, so a standardized IAM policy and kubeconfig-delivery pattern across NCPs are the goal. Identity aligned with the Software Reference Guide user personas is provided here as well.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                  | Keyword    | NCP (Operator) | ISP (Tenant) |
| :----- | :------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :----------- |
| KUB-21 | After installation, provide ISP access to the Kubernetes API/kubectl, gated by an IAM policy that requires authentication and authorization. | **MUST**   | Provides       | Consumes     |
| KUB-22 | Deliver interaction with the deployed cluster through an exposed AI Server.                                                                  | **SHOULD** | Provides       | Consumes     |

### Topology-Aware Placement and NUMA Alignment

Placement and scheduling policy belongs to the ISP; the NCP's job is to expose the cluster topology and enable the node-level controls the ISP's scheduler needs. Neither the GPU Operator nor the Network Operator makes placement or NUMA decisions; they provision node-level GPU and networking capabilities and publish the labels and resources a scheduler reads. The NCP exposes topology as Kubernetes node labels, with [Topograph](https://github.com/NVIDIA/topograph) producing the fabric labels (NVLink Domain, NVSwitch fabric, rack, and IB partition), and [Node Feature Discovery](https://github.com/kubernetes-sigs/node-feature-discovery) and [NVIDIA GPU Feature Discovery](https://github.com/NVIDIA/gpu-feature-discovery) supplying the complementary hardware labels. It enables NUMA alignment through the kubelet Topology Manager, CPU Manager, and Memory Manager, plus device NUMA affinity from the NVIDIA device plugin, and supports Dynamic Resource Allocation and CDI so GPUs and NICs are co-allocated with topology awareness.

Given that substrate, the ISP scheduler (for example, [KAI Scheduler](https://github.com/NVIDIA/KAI-Scheduler), Volcano, or Kueue) enforces the placement policy, and the kubelet Topology Manager enforces NUMA pinning on the nodes the scheduler selects.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                                                                                                             | Keyword    | NCP (Operator)                   | ISP (Tenant) |
| :----- | :------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------ | :--------- | :------------------------------- | :----------- |
| KUB-23 | Expose cluster topology as node labels (NVLink Domain, NVSwitch fabric ID, rack/IB partition, region/zone), for example using [Topograph](https://github.com/NVIDIA/topograph) with [Node Feature Discovery](https://github.com/kubernetes-sigs/node-feature-discovery) and [NVIDIA Feature Discovery](https://github.com/NVIDIA/gpu-feature-discovery) | **MUST**   | Owns                             | Consumes     |
| KUB-24 | Enable NUMA alignment: kubelet Topology Manager (single-numa-node), CPU Manager (static), Memory Manager, and device NUMA affinity through the NVIDIA device plugin.                                                                                                                                                                                    | **MUST**   | Owns                             | Consumes     |
| KUB-25 | Support [Dynamic Resource Allocation](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/dra-intro-install.html) (DRA) and CDI for topology-aware GPU-NIC co-allocation. Placement and scheduling policy (TP within NVLink Domain; EP/PP within rack/IB partition; NUMA pinning; replica and tenant anti-affinity) is ISP-owned.       | **SHOULD** | Owns substrate; exposes topology | Owns policy  |

### Additional Kubernetes Components

This layer runs on top of the managed cluster, and the platform, meaning the ISP, owns most of it. It spans platform packaging and networking ([Helm](https://helm.sh/), [Cilium](https://cilium.io/), [Multus](https://github.com/k8snetworkplumbingwg/multus-cni)), scheduling ([KAI Scheduler](https://github.com/NVIDIA/KAI-Scheduler), [Volcano](https://volcano.sh/), or [Kueue](https://kueue.sigs.k8s.io/) with a custom extender), node lifecycle and OS imaging ([Cluster API](https://cluster-api.sigs.k8s.io/) over [NVIDIA Infra Controller](https://docs.nvidia.com/infra-controller/documentation/home), [Tinkerbell](https://tinkerbell.org/), or NCP-native bare-metal APIs), GitOps and progressive delivery ([Argo](https://argoproj.github.io/)), in-region image distribution ([Harbor](https://goharbor.io/) with an [NGC](https://catalog.ngc.nvidia.com/) pull-through cache, plus [Spegel](https://github.com/spegel-org/spegel) or [Dragonfly](https://d7y.io/) for peer-to-peer distribution), the driver, CUDA, NCCL, and fabric-manager lifecycle (pinned per node pool through the [GPU Operator](https://docs.nvidia.com/datacenter/cloud-native/gpu-operator/latest/index.html)), health remediation, and model-weight distribution (an object store, a GPUDirect-Storage warm cache, a per-node NVMe cache, and peer-to-peer weight pull over the east-west RDMA fabric).

Because the platform owns most of these, the table below carries only the NCP responsibilities in this layer: enable Dynamic Resource Allocation, own firmware, BIOS, and BMC, provide the node base images, provision and join nodes, expose the managed-Kubernetes and BMaaS APIs that the platform's health remediation calls, and provide the GPUDirect-Storage filesystem and east-west RDMA fabric that model-weight distribution depends on. The remaining rows are marked as INFO.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                              | Keyword    | NCP (Operator)             | ISP (Tenant)         |
| :----- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------------------- | :------------------- |
| KUB-26 | Helm for packaging/rollout; Cilium CNI (eBPF, NetworkPolicy, Hubble) with BlueField DPU offload; Multus for RDMA secondary interfaces.                                                                                                                                   | **INFO**   | Substrate (DPU offload)    | Owns                 |
| KUB-27 | KAI/Volcano/Kueue + custom scheduler extender (gang scheduling, topology-aware placement, priority preemption, quotas, MIG-aware bin packing).                                                                                                                           | **INFO**   | n/a                        | Owns                 |
| KUB-28 | Cluster API (per-NCP providers); NVIDIA Infra Controller / Tinkerbell / NCP-native bare-metal APIs for OS imaging; Grove for [Dynamo](https://github.com/ai-dynamo/dynamo) using custom resources.                                                                       | **INFO**   | Provides (bare-metal APIs) | Drives (Cluster API) |
| KUB-29 | Argo CD + Rollouts + Workflows for GitOps and progressive delivery; [cert-manager](https://cert-manager.io/), [external-dns](https://github.com/kubernetes-sigs/external-dns), [external-secrets](https://external-secrets.io/) ([Vault](https://www.vaultproject.io/)). | **INFO**   | n/a                        | Owns                 |
| KUB-30 | Harbor in-region registry with NGC pull-through cache and [Cosign](https://github.com/sigstore/cosign) signature verification; Spegel/Dragonfly for P2P image distribution.                                                                                              | **INFO**   | n/a                        | Owns                 |
| KUB-31 | Driver/CUDA/NCCL/fabric-manager pinned per node pool through the GPU Operator; canary through staging with rollback on regression.                                                                                                                                       | **INFO**   | Provides (base images)     | Owns (pins/canaries) |
| KUB-32 | Validated combination matrix (driver × CUDA × NCCL × engine) maintained by an automated test suite.                                                                                                                                                                      | **INFO**   | n/a                        | Owns                 |
| KUB-33 | Multiple driver streams run simultaneously across pools (H100 production + B200 latest-features) without forced lockstep migration.                                                                                                                                      | **INFO**   | Provides (pools)           | Owns                 |
| KUB-34 | Node-pool lifecycle through the managed-K8s API + Cluster API; the platform declares pools/surge/topology, and NCP provisions metal and joins nodes.                                                                                                                     | **INFO**   | Provisions metal           | Declares pools       |
| KUB-35 | Firmware/BIOS/BMC management remains the NCP's responsibility; burn-in pipeline gates promotion into the production pool using a staging taint and the acceptance tests.                                                                                                 | **MUST**   | Owns                       | n/a                  |
| KUB-36 | Health-remediation controller monitors DCGM, node-problem-detector, and kernel logs; cordons/drains autonomously; triggers NCP-side reboot/RMA through managed K8s and BMaaS APIs.                                                                                       | **INFO**   | Executes (reboot/RMA)      | Owns (controller)    |
| KUB-37 | Model weights: object store canonical; per-region warm cache on GPUDirect-Storage-enabled shared FS; per-node NVMe cache with LRU eviction and pinning.                                                                                                                  | **INFO**   | Provides (GDS FS)          | Owns (caching)       |
| KUB-38 | Pull weights over the back-end RDMA fabric fast enough to hydrate a cold node (\~400 GB 405B FP8) in under 60 s; chunking/P2P puller with GPUDirect RDMA into HBM is a platform choice.                                                                                  | **SHOULD** | Provides (RDMA fabric)     | Owns (puller)        |

## Inference Routing Data Plane

The inference routing requirements in this section are covered in greater detail in the [edge connectivity reference architecture document](https://docs.google.com/document/d/1VCzBUZ9_lX8nvmsvRQC_kaPBEOWdmujEj_sC97_Gc3A/edit?usp=sharing).

Comprehensive recommendations for inference routing are covered in the [Inference Reference Architecture](https://docs.nvidia.com/ncx/ncp-inference-ra/latest/index.html) document.

**Requirements Summary**

| Req ID | Requirement                                                                                                    | Keyword  | NCP (Operator) | ISP (Tenant)   |
| :----- | :------------------------------------------------------------------------------------------------------------- | :------- | :------------- | :------------- |
| RT-1   | Inference routing details are documented in the edge-connectivity RA and the Inference Reference Architecture. | **INFO** | Sizes (edge)   | Owns (routing) |

## Autoscaling and Scheduling

Autoscaling is where the NCP and the ISP intersect: the ISP's platform drives Kubernetes autoscaling, and the NCP supplies and reclaims the machines underneath. It operates at two scales: replica scaling within a deployment, which the ISP owns, and fleet scaling of the underlying pools, which depends on the NCP. On contraction, scale-down cordons a node, drains it, and removes it from the cluster. The fleet autoscaler tracks free GPU-seconds per pool, SKU, and region; bursts on demand from the NCP when a reserved pool falls below its headroom target; drains the spot pool first; and communicates with the Cluster API to the NCP's machine controllers. The fleet-facing items are the NCP's; the replica-scaling behavior is ISP-owned.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                                                | Keyword    | NCP (Operator)            | ISP (Tenant)   |
| :----- | :----------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------------------ | :------------- |
| AS-1   | Scaling down removes the node from the cluster and auto-cordons nodes slated for removal.                                                                                                                                                                                                  | **SHOULD** | Executes (node removal)   | Triggers       |
| AS-2   | Replica-autoscaler signals: token-queue depth, TTFT p95 vs SLO, inter-token latency p95 vs SLO, KV-cache occupancy, active requests, 1–5 min predictive look-ahead.                                                                                                                        | **INFO**   | n/a                       | Owns           |
| AS-3   | Behaviors: scale up aggressively / down conservatively (5–15 min cooldown); scale-to-zero with snapshot/restore; warm pools; pre-warm; cold-start budgets \<30 s (7B/8B), \<90 s (70B), \<180 s (400B+) using pre-pulled images, pre-cached weights, [CRIU](https://criu.org/), lazy load. | **INFO**   | Provides (substrate)      | Owns           |
| AS-4   | Fleet autoscaler: free GPU-seconds per pool × SKU × region; on-demand NCP burst below headroom; spot-pool drain-first; Cluster API to NCP machine controllers.                                                                                                                             | **INFO**   | Provides (burst capacity) | Owns (decides) |
| AS-5   | Scheduling (Volcano/Kueue): bin-pack within SLO; gang scheduling (TP ≥2, PP, EP for MoE); TP within NVLink Domain, PP/EP within rack/IB partition, never split TP across NUMA; preempt low-priority batch on 30 s notice with checkpointing.                                               | **INFO**   | Exposes (topology)        | Owns           |

## Telemetry and Observability

The NCP exports the telemetry the ISP needs to operate the cluster and meet its SLAs, applying the [AI Cloud Requirements telemetry bar](https://docs.nvidia.com/dsx/ncp/nvidia-requirements-for-ai-clouds/telemetry-requirements) and the [Software Reference Guide, Telemetry and Observability](https://docs.nvidia.com/dsx/ncp/part-1-software-reference-guide/telemetry-and-observability) to GB300 NVL72. NVIDIA provides the building blocks through DSX OS components: [DCGM](https://github.com/NVIDIA/dcgm-exporter) for GPU health and metrics, [Fleet Intelligence](https://docs.nvidia.com/fleet-intelligence/latest/index.html) for fleet-wide health and visibility, and [NVSentinel](https://docs.nvidia.com/nvsentinel/getting-started/overview/) for fault detection and remediation.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                           | Keyword    | NCP (Operator) | ISP (Tenant) |
| :----- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :--------- | :------------- | :----------- |
| TEL-1  | Use the DSX OS telemetry components: [DCGM](https://github.com/NVIDIA/dcgm-exporter), [Fleet Intelligence](https://docs.nvidia.com/fleet-intelligence/latest/index.html), [NVSentinel.](https://docs.nvidia.com/nvsentinel/getting-started/overview/) | **SHOULD** | Owns / Exports | Consumes     |

## Identity Management

The NCP should provide identity aligned with the [Software Reference Guide user personas](https://docs.nvidia.com/dsx/guides/ncp-software-reference-guide/introduction#user-personas). Obtaining kubectl and the associated identity policies is covered in the [Exposing Kubernetes Control Plane - Kubeconfig Acquisition](#exposing-kubernetes-control-plane-kubeconfig-acquisition) section.

**Requirements Summary**

| Req ID | Requirement                                          | Keyword    | NCP (Operator) | ISP (Tenant) |
| :----- | :--------------------------------------------------- | :--------- | :------------- | :----------- |
| IDM-1  | Provide identity aligned with the SRG user personas. | **SHOULD** | Provides       | Consumes     |

## Operations

Operational requirements apply the [AI Cloud Requirements Service Delivery SLAs](https://docs.nvidia.com/dsx/ncp/nvidia-requirements-for-ai-clouds/service-delivery-sl-as) to GB300 NVL72. The NCP staffs the cluster with a dedicated technical specialist reachable by NVIDIA, runs a monitored Slack channel or equivalent, and provides 24x7 support under its standard incident-severity procedures. It communicates service-impacting incidents and both planned and unplanned maintenance to NVIDIA, and it allows NVIDIA to schedule planned maintenance windows using APIs or console tools. The NCP also remediates critical vulnerabilities promptly and discloses them transparently, and it runs an incident toolchain built on [PagerDuty](https://www.pagerduty.com/), [Statuspage](https://www.atlassian.com/software/statuspage), a war-room template, and a customer-communications playbook.

NVIDIA-stack-specific operational failure modes the platform must handle automatically:

* DCGM continuously monitors XIDs (driver-level GPU fault codes), and the appropriate response depends on the fault reported. When DCGM detects an XID, the general remediation sequence is to cordon the node, drain its workloads, and raise an RMA ticket. However, the urgency varies by fault type: XID 79 (GPU falling off the bus) and XID 48 (double-bit ECC errors) both require immediate cordoning, while XID 31 and XID 13 (MMU faults) permit a softer response.
* Double-bit ECC errors warrant immediate node cordoning regardless of workload state.
* NCCL hangs, and collective timeouts should be handled by killing the affected replica and attempting a restart. If the same issue repeats more than N times, escalate to on-call. In parallel, check the fabric manager state and verify SHARP tree health, since either can be the root cause.
* InfiniBand and RoCE link flaps call for quarantining the affected node and opening a fabric ticket. UFM events should be pulled and correlated to determine whether the flap is isolated or part of a broader fabric issue.
* If the Fabric Manager service crashes (particularly on NVL72 systems), the first response is automatic restart. If the crashes persist, drain the entire rack rather than individual nodes.
* NVSwitch errors require cordoning off the affected node(s). If rack-scale isolation is needed, drain workloads to non-NVL72 capacity.
* Thermal throttling should trigger load shedding on the affected node. Cooling investigation should be opened with the Network and Compute Platform (NCP) team.
* Silent data corruption is addressed proactively: idle replicas run periodic checksum self-tests, and any node that fails or raises suspicion is cordoned immediately.
* If a node refresh results in a driver or CUDA version mismatch, the admission controller will automatically block scheduling onto that node until the version discrepancy is resolved.

Change management:

* All platform changes flow through GitOps using Argo, so no manual configuration changes reach production outside a reviewed and committed change set.
* New features and model updates are delivered progressively. A rollout starts at canary scale, then moves through 1%, 10%, and 50% of traffic before reaching full deployment. At each stage, SLO metrics are evaluated automatically, and the system rolls back without human intervention if any regression is detected.
* Feature flags are used to isolate risk during launches, allowing new capabilities to be enabled or disabled independently of a deployment boundary.
* A per-model evaluation suite gates model and engine version upgrades. A new version cannot proceed to production if it introduces a quality regression on the golden datasets for that model.
* Driver, CUDA, and NCCL upgrades follow a separate gate: the new version must complete a soak period in the staging pool and pass a benchmark regression check before it is eligible for production rollout.

**Requirements Summary**

| Req ID | Requirement                                                                                                                                                                                                                                                                                                                                                                                                                                                 | Keyword  | NCP (Operator)               | ISP (Tenant)               |
| :----- | :---------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------------- | :------- | :--------------------------- | :------------------------- |
| OPS-1  | Dedicated technical specialist/engineer available to NVIDIA.                                                                                                                                                                                                                                                                                                                                                                                                | **MUST** | Owns                         | Consumes                   |
| OPS-2  | Monitored Slack (or equivalent) channel.                                                                                                                                                                                                                                                                                                                                                                                                                    | **MUST** | Owns                         | Consumes                   |
| OPS-3  | 24x7 support per partner-standard incident-severity procedures.                                                                                                                                                                                                                                                                                                                                                                                             | **MUST** | Owns                         | Consumes                   |
| OPS-4  | Communicate service-impacting maintenance events, planned and unplanned, to NVIDIA.                                                                                                                                                                                                                                                                                                                                                                         | **MUST** | Owns                         | Consumes                   |
| OPS-5  | Provide maintenance window scheduling to NVIDIA using API/console tools.                                                                                                                                                                                                                                                                                                                                                                                    | **MUST** | Provides                     | Consumes                   |
| OPS-6  | Remediate critical vulnerabilities promptly and disclose them transparently.                                                                                                                                                                                                                                                                                                                                                                                | **MUST** | Owns                         | Consumes                   |
| OPS-7  | Maintain an Incident Management + Status Page + War Room + Customer Comms playbook.                                                                                                                                                                                                                                                                                                                                                                         | **MUST** | Owns                         | Consumes                   |
| OPS-8  | The platform auto-handles NVIDIA-stack failure modes: DCGM XIDs (79 immediate, 31/13 soft, 48 immediate) → cordon/drain/RMA; double-bit ECC → cordon; NCCL hangs → kill/restart/escalate, check fabric-manager + SHARP; IB/RoCE flaps → quarantine, UFM correlate; Fabric Manager crash → restart, drain rack if persistent; NVSwitch errors → cordon; thermal → load shed; SDC → checksum self-tests; driver/CUDA mismatch → admission controller rejects. | **MUST** | Executes (RMA/reboot/fabric) | Owns (detect/cordon/drain) |
| OPS-9  | Change management: all platform changes through GitOps (Argo); progressive delivery canary→1→10→50→100% with auto-rollback on SLO regression; feature flags; model/engine rollouts gated by per-model eval suite; driver/CUDA/NCCL rollouts gated by staging-pool soak + benchmark regression.                                                                                                                                                              | **INFO** | n/a                          | Owns                       |

## SLOs and Non-Functional Requirements

Published, contractually backed where applicable.

SLO requirements are outlined in the [Requirements for AI Clouds Service Delivery SLAs](https://docs.nvidia.com/dsx/guides/nvidia-requirements-for-ai-clouds/service-delivery-sl-as).

**Requirements Summary**

| SLI                             | Serverless target | Dedicated target |
| :------------------------------ | :---------------- | :--------------- |
| API availability (HTTP 2xx/5xx) | 99.9% / month     | 99.95% / month   |
| Meter accuracy                  | ±0.5% / month     | ±0.1%            |
| Support response (Sev-1)        | 30 min            | 15 min           |