Glossary for NIM Metadata API#

Use this glossary to understand key terms used throughout the NIM Metadata API documentation.

G#

Browse terms starting with G.

GPU-agnostic profile

A profile that is not tied to a particular GPU and runs anywhere with sufficient vRAM. These profiles omit gpu and gpu_device entirely and can carry an empty tested_gpu_devices array.

Related: Core Concepts Overview

gpu_device

The PCI identifier of the GPU a profile targets, in the form pciDeviceId:pciVendorId. This value is precise and machine-comparable, so it is the field to match on when you evaluate hardware fit.

2330:10de

Related: Response Fields

I#

Browse terms starting with I.

Inference backend

The serving backend a profile uses, such as tensorrt_llm, vllm, or sglang. Treat this value as an open string rather than an enumeration, because new values appear without notice.

Related: Notes for Automated Agents

M#

Browse terms starting with M.

Metadata coverage

The extent to which published NIM microservices have API-backed metadata. Metadata is attached at publish time, so coverage grows as NIM microservices are released and re-released.

Related: Current Limitations

N#

Browse terms starting with N.

NIM microservice

An NVIDIA NIM microservice, which is a GPU-accelerated container that packages a model, the inference stack, and a unified API. Each NIM microservice ships one or more profiles.

Related: NIM Metadata API Documentation

P#

Browse terms starting with P.

Pipeline parallel size

The number of pipeline stages a profile splits the model across, recorded as pipeline_parallel_size. It combines with tensor parallel size to determine the total number of GPUs a profile requires.

Related: Choose a Profile for Your Hardware

Precision

The numeric precision a profile uses, such as fp8, bf16, or fp16. Lower precision generally favors throughput, and higher precision generally favors accuracy.

Related: Response Fields

Profile

A deployable configuration of a NIM microservice that pairs the model with a GPU target, a numeric precision, an inference backend, and a parallelism layout. A single NIM microservice can publish a handful of profiles or several hundred.

Related: Core Concepts Overview

Profile identifier

The unique identifier for a profile, recorded as profile_id, and the value you supply at deployment time. Treat it as opaque: compare it for equality, and never parse it or derive meaning from its structure.

Related: Current Limitations

R#

Browse terms starting with R.

Resource identifier

The slash-separated repository path that identifies a NIM microservice container, recorded as resourceId. It takes the form {org}/{repo-name} or {org}/{team-name}/{repo-name}, and it must not contain a colon.

nim/meta/llama-3.1-8b-instruct

Related: Prerequisites

T#

Browse terms starting with T.

Tag

The container tag, exactly as it appears in the catalog, such as 1.2.0 or latest. Metadata is stored per tag, so the presence of metadata for one tag says nothing about another.

Related: Endpoints

Tensor parallel size

The number of GPUs a profile shards the model across tensor-wise, recorded as tensor_parallel_size.

Related: Choose a Profile for Your Hardware

Tested GPU devices

The PCI identifiers a profile has been validated on, recorded as tested_gpu_devices. An empty array means the profile was explicitly tested on nothing, and an absent field means no testing information was recorded.

Related: Core Concepts Overview

Throughput

A recorded throughput figure for a profile. Units and measurement conditions are not standardized, so figures are not comparable across NIM microservices.

Related: Current Limitations

V#

Browse terms starting with V.

vRAM floor

The minimum vRAM a profile requires per GPU, recorded as min_vram_per_device_gb. The value is per device rather than a total across devices, and it is the primary hardware-fit field.

Related: Response Fields

W#

Browse terms starting with W.

Workspace hash

A digest over the workspace contents of a profile, recorded as workspace_hash. Use it to verify supply-chain integrity after download. This field is optional today, so do not assume its presence.

Related: Current Limitations