Core Concepts Overview for NIM Metadata API#
To interpret a response from the NIM Metadata API correctly, you first need to understand profiles, how the API expresses GPU targeting, and the serialization rules that govern which fields appear at all. For the endpoint and field listings, refer to the Reference Overview.
Profiles#
A profile is a deployable configuration of a NIM microservice. Each profile pairs the model with a GPU target, a numeric precision, an inference backend, and a parallelism layout. A single NIM microservice can publish a handful of profiles or several hundred.
Profiles are the section of the response you use most. The profile you select
determines the hardware you need, so comparing profiles is the primary reason
to call this API. Every profile carries a profile_id, which is the value you
supply at deployment time.
Treat profile_id as an opaque identifier. Compare it for equality, and never
parse it or derive meaning from its structure. Identifiers generated by
different versions of the publishing toolchain were computed differently, so the
internal structure is not stable across publication vintages.
Hardware Fit#
The primary hardware-fit field is min_vram_per_device_gb, which states the
minimum vRAM required per GPU rather than a total across GPUs. A profile with
tensor_parallel_size set to 4 and min_vram_per_device_gb set to 80.0
requires four GPUs with 80 GB each, not 80 GB in aggregate.
The number of GPUs a profile requires is the product of its tensor parallel size and its pipeline parallel size, treating an absent value as 1:
required_gpus = (tensor_parallel_size or 1) × (pipeline_parallel_size or 1)
GPU Targeting#
Not every profile targets a specific GPU. Many NIM microservices ship profiles
that run anywhere with sufficient vRAM. Those profiles omit gpu and
gpu_device entirely, and can carry an empty tested_gpu_devices array.
An empty or absent GPU field is not a broken response. It means the profile is not tied to a particular GPU. Read it as “no GPU restriction recorded” rather than “no GPUs supported.”
The API preserves this distinction precisely, as described in the following table:
What You See |
What It Means |
|---|---|
|
This profile targets that specific GPU. |
|
Validated on these specific devices. |
|
No validation information recorded. GPUs were not tested with the profile. |
|
No testing information recorded. |
Comparing gpu and gpu_device#
Both fields describe the target GPU, but they are not interchangeable:
gpu_deviceis a PCI identifier in the formpciDeviceId:pciVendorId, such as2330:10de. It is precise and machine-comparable, so match on this field.gpuis a human-readable name such asH100orA100_1x. Formats vary between NIM microservices, so it is not a reliable join key. Display it, but do not match on it.
When both fields are present, they describe the same GPU. When only gpu is
present, you can show it to a reader but you cannot reliably match it
programmatically, so treat that profile as unverified rather than assuming a
match.
Absent Means Unknown#
The service serializes responses with non-null inclusion, which means it never
emits null. If a field has no value, the key is absent entirely. A response
contains what was supplied at publish time and nothing else, so two profiles in
the same response body can carry different sets of keys.
The consequence matters for every decision you make from this data. An absent
field records the absence of information, not a negative value. A missing
feat_lora does not mean LoRA adapters are unsupported; it means nothing was
recorded. Never treat an absent boolean as false, or an absent number as zero.
There is one deliberate exception to the omission rule. An explicitly empty
tested_gpu_devices array is preserved, so “tested on nothing” stays
distinguishable from “unknown.”
For the complete set of serialization rules, refer to Response Fields.
Schema Evolution#
The set of fields grows as the schema evolves, so a response can contain keys that postdate your integration. Parse permissively, ignore keys you do not recognize, and never fail on an unrecognized key.
Treat precision and inference_backend as open strings rather than
enumerations. New values appear without notice, so carry an unrecognized value
through instead of rejecting it. For the full set of rules that keep an
integration working across schema changes, refer to
Notes for Automated Agents.