Glossary for NIM Metadata API#
Use this glossary to understand key terms used throughout the NIM Metadata API documentation.
G#
Browse terms starting with G.
- GPU-agnostic profile
A profile that is not tied to a particular GPU and runs anywhere with sufficient vRAM. These profiles omit
gpuandgpu_deviceentirely and can carry an emptytested_gpu_devicesarray.Related: Core Concepts Overview
- gpu_device
The PCI identifier of the GPU a profile targets, in the form
pciDeviceId:pciVendorId. This value is precise and machine-comparable, so it is the field to match on when you evaluate hardware fit.2330:10de
Related: Response Fields
I#
Browse terms starting with I.
- Inference backend
The serving backend a profile uses, such as
tensorrt_llm,vllm, orsglang. Treat this value as an open string rather than an enumeration, because new values appear without notice.Related: Notes for Automated Agents
M#
Browse terms starting with M.
- Metadata coverage
The extent to which published NIM microservices have API-backed metadata. Metadata is attached at publish time, so coverage grows as NIM microservices are released and re-released.
Related: Current Limitations
N#
Browse terms starting with N.
- NIM microservice
An NVIDIA NIM microservice, which is a GPU-accelerated container that packages a model, the inference stack, and a unified API. Each NIM microservice ships one or more profiles.
Related: NIM Metadata API Documentation
P#
Browse terms starting with P.
- Pipeline parallel size
The number of pipeline stages a profile splits the model across, recorded as
pipeline_parallel_size. It combines with tensor parallel size to determine the total number of GPUs a profile requires.Related: Choose a Profile for Your Hardware
- Precision
The numeric precision a profile uses, such as
fp8,bf16, orfp16. Lower precision generally favors throughput, and higher precision generally favors accuracy.Related: Response Fields
- Profile
A deployable configuration of a NIM microservice that pairs the model with a GPU target, a numeric precision, an inference backend, and a parallelism layout. A single NIM microservice can publish a handful of profiles or several hundred.
Related: Core Concepts Overview
- Profile identifier
The unique identifier for a profile, recorded as
profile_id, and the value you supply at deployment time. Treat it as opaque: compare it for equality, and never parse it or derive meaning from its structure.Related: Current Limitations
R#
Browse terms starting with R.
- Resource identifier
The slash-separated repository path that identifies a NIM microservice container, recorded as
resourceId. It takes the form{org}/{repo-name}or{org}/{team-name}/{repo-name}, and it must not contain a colon.nim/meta/llama-3.1-8b-instruct
Related: Prerequisites
T#
Browse terms starting with T.
- Tag
The container tag, exactly as it appears in the catalog, such as
1.2.0orlatest. Metadata is stored per tag, so the presence of metadata for one tag says nothing about another.Related: Endpoints
- Tensor parallel size
The number of GPUs a profile shards the model across tensor-wise, recorded as
tensor_parallel_size.Related: Choose a Profile for Your Hardware
- Tested GPU devices
The PCI identifiers a profile has been validated on, recorded as
tested_gpu_devices. An empty array means the profile was explicitly tested on nothing, and an absent field means no testing information was recorded.Related: Core Concepts Overview
- Throughput
A recorded throughput figure for a profile. Units and measurement conditions are not standardized, so figures are not comparable across NIM microservices.
Related: Current Limitations
V#
Browse terms starting with V.
- vRAM floor
The minimum vRAM a profile requires per GPU, recorded as
min_vram_per_device_gb. The value is per device rather than a total across devices, and it is the primary hardware-fit field.Related: Response Fields
W#
Browse terms starting with W.
- Workspace hash
A digest over the workspace contents of a profile, recorded as
workspace_hash. Use it to verify supply-chain integrity after download. This field is optional today, so do not assume its presence.Related: Current Limitations