Core Concepts Overview for NIM Metadata API#

To interpret a response from the NIM Metadata API correctly, you first need to understand profiles, how the API expresses GPU targeting, and the serialization rules that govern which fields appear at all. For the endpoint and field listings, refer to the Reference Overview.

Profiles#

A profile is a deployable configuration of a NIM microservice. Each profile pairs the model with a GPU target, a numeric precision, an inference backend, and a parallelism layout. A single NIM microservice can publish a handful of profiles or several hundred.

Profiles are the section of the response you use most. The profile you select determines the hardware you need, so comparing profiles is the primary reason to call this API. Every profile carries a profile_id, which is the value you supply at deployment time.

Treat profile_id as an opaque identifier. Compare it for equality, and never parse it or derive meaning from its structure. Identifiers generated by different versions of the publishing toolchain were computed differently, so the internal structure is not stable across publication vintages.

Hardware Fit#

The primary hardware-fit field is min_vram_per_device_gb, which states the minimum vRAM required per GPU rather than a total across GPUs. A profile with tensor_parallel_size set to 4 and min_vram_per_device_gb set to 80.0 requires four GPUs with 80 GB each, not 80 GB in aggregate.

The number of GPUs a profile requires is the product of its tensor parallel size and its pipeline parallel size, treating an absent value as 1:

required_gpus = (tensor_parallel_size or 1) × (pipeline_parallel_size or 1)

GPU Targeting#

Not every profile targets a specific GPU. Many NIM microservices ship profiles that run anywhere with sufficient vRAM. Those profiles omit gpu and gpu_device entirely, and can carry an empty tested_gpu_devices array.

An empty or absent GPU field is not a broken response. It means the profile is not tied to a particular GPU. Read it as “no GPU restriction recorded” rather than “no GPUs supported.”

The API preserves this distinction precisely, as described in the following table:

What You See

What It Means

gpu_device present

This profile targets that specific GPU.

tested_gpu_devices with one or more entries

Validated on these specific devices.

tested_gpu_devices present and empty

No validation information recorded. GPUs were not tested with the profile.

tested_gpu_devices absent

No testing information recorded.

Comparing gpu and gpu_device#

Both fields describe the target GPU, but they are not interchangeable:

  • gpu_device is a PCI identifier in the form pciDeviceId:pciVendorId, such as 2330:10de. It is precise and machine-comparable, so match on this field.

  • gpu is a human-readable name such as H100 or A100_1x. Formats vary between NIM microservices, so it is not a reliable join key. Display it, but do not match on it.

When both fields are present, they describe the same GPU. When only gpu is present, you can show it to a reader but you cannot reliably match it programmatically, so treat that profile as unverified rather than assuming a match.

Absent Means Unknown#

The service serializes responses with non-null inclusion, which means it never emits null. If a field has no value, the key is absent entirely. A response contains what was supplied at publish time and nothing else, so two profiles in the same response body can carry different sets of keys.

The consequence matters for every decision you make from this data. An absent field records the absence of information, not a negative value. A missing feat_lora does not mean LoRA adapters are unsupported; it means nothing was recorded. Never treat an absent boolean as false, or an absent number as zero.

There is one deliberate exception to the omission rule. An explicitly empty tested_gpu_devices array is preserved, so “tested on nothing” stays distinguishable from “unknown.”

For the complete set of serialization rules, refer to Response Fields.

Schema Evolution#

The set of fields grows as the schema evolves, so a response can contain keys that postdate your integration. Parse permissively, ignore keys you do not recognize, and never fail on an unrecognized key.

Treat precision and inference_backend as open strings rather than enumerations. New values appear without notice, so carry an unrecognized value through instead of rejecting it. For the full set of rules that keep an integration working across schema changes, refer to Notes for Automated Agents.