Response Fields for NIM Metadata API#

The response body has four top-level sections, and all four are optional. Real responses contain a subset of the fields described on this page. A complete example with every field populated is available in the API reference documentation.

The following table describes the top-level sections:

Section

Type

Description

profiles

array of object

The deployable configurations. This is the section you use most.

release_details

object

Release-level facts, such as version, documentation, compliance, and supported architectures.

features

object

Capability flags, such as context window, tool calling, and backend support.

related_urls

array of string

Links to related resources for this NIM microservice.

Serialization Rules#

The following rules hold across the entire JSON response body and are enforced by the service. Use these rules to interpret anything you get back.

  1. The service never emits null. The service serializes with non-null inclusion. If a field has no value, the key is absent entirely.

  2. Absent means unknown, not false. A missing feat_lora does not mean LoRA adapters are unsupported; it means nothing was recorded. Never treat an absent boolean as false, or an absent number as 0.

  3. Only populated fields appear. A response contains what was supplied at publish time and nothing else. Two profiles in the same response body can carry different sets of keys and values.

  4. One deliberate exception applies. An explicitly empty tested_gpu_devices array is preserved, so “tested on nothing” stays distinguishable from “unknown.”

  5. Tolerate keys you do not recognize. The set of fields grows as the schema evolves, so a response body can contain keys that postdate your integration. Parse permissively, ignore what you do not know, and never fail on an unrecognized key.

Profiles#

The profiles array holds the deployable configurations. Only profile_id and tested_gpu_devices are guaranteed present, so treat every other field as optional.

The following table describes the profile fields:

Field

Type

Description

profile_id

string

Unique identifier for this profile. Use this value when selecting a profile at deployment time. Always present.

gpu

string

Human-readable GPU this profile targets, such as H100 or A100. Absent on GPU-agnostic profiles.

gpu_device

string

PCI identifier of the target GPU, in the form pciDeviceId:pciVendorId, such as 2330:10de. More precise than gpu. Absent on GPU-agnostic profiles.

tested_gpu_devices

array of string

PCI identifiers this profile has been validated on. Can be empty, as described in Core Concepts Overview.

precision

string

Numeric precision, such as fp8, bf16, or fp16.

inference_backend

string

Serving backend, such as tensorrt_llm, vllm, or sglang.

tensor_parallel_size

integer

Number of GPUs the model is sharded across tensor-wise.

pipeline_parallel_size

integer

Number of pipeline stages.

min_vram_per_device_gb

number

Minimum vRAM required per GPU, not a total across GPUs. This is the primary hardware-fit field. A profile with tensor_parallel_size set to 4 and min_vram_per_device_gb set to 80.0 requires four GPUs with 80 GB each, not 80 GB in aggregate.

disk_space_per_device_gb

number

Disk space required per device, in GB.

system_memory_per_device_gb

number

Host RAM required per device, in GB.

min_cpu_count

integer

Minimum CPU cores.

min_driver_version

string

Minimum NVIDIA driver version.

max_gpu_count

integer

Maximum GPUs this profile can use.

form_factor

string

Physical GPU form factor, such as SXM or PCIe.

workspace_hash

string

Digest over the workspace contents of the profile. Use it to verify supply-chain integrity after download.

throughput

integer

Recorded throughput figure. Units are not standardized, as described in Current Limitations.

feat_lora

boolean

Whether this profile supports LoRA adapters.

buildable_profile

boolean

Whether the engine for the profile can be built locally rather than downloaded prebuilt.

languages

array of string

Language codes the profile supports, such as ["en-US", "ja-JP"]. Typically present only on speech and translation NIM microservices.

Release Details#

The release_details object holds facts about the release of a container. All fields are optional.

The following table describes the release fields:

Field

Type

Description

id

string

Release identifier.

container_version

string

Version of the container this metadata describes.

model_card

string

Model card content, as Markdown.

minimum_gpu_supported

string

Lowest GPU generation supported by the release as a whole.

compliance_certifications

array of string

Certifications held, such as ["ISO", "SOC2"].

documentation_url

string

Link to the documentation for the NIM microservice.

architecture

array of string

CPU architectures supported, such as ["AMD 64", "ARM 64"].

is_gov_ready

boolean

Whether the release is approved for government deployment.

allowed_deployment_regions

array of string

Regions the release can be deployed in, such as ["us-west-1"]. An absent value means no recorded restriction. Confirm license terms separately rather than inferring them from absence.

Features#

The features object holds capability flags for the NIM microservice. All fields are optional and boolean unless noted otherwise.

The following table describes the feature flags:

Field

Type

Description

context_window

integer

Maximum context length in tokens.

reasoning_support

boolean

Supports reasoning-style generation.

tool_calling

boolean

Supports tool or function calling.

parallel_tool_calling

boolean

Supports multiple simultaneous tool calls.

vgpu_support

boolean

Runs on virtualized GPUs.

trt_llm_buildable_support

boolean

TensorRT-LLM engines can be built locally.

prebuilt_engines

boolean

Ships prebuilt inference engines.

vllm_support

boolean

Supports the vLLM backend.

sglang_support

boolean

Supports the SGLang backend.

trtllm_pytorch_support

boolean

Supports the TensorRT-LLM PyTorch backend.

suffix_support

boolean

Supports suffix or fill-in-the-middle completion.

boost_decode

boolean

Boosted decoding optimization available.

steer_decode

boolean

Steered decoding available.

rule_decode

boolean

Rule-constrained decoding available.

serve_scale

boolean

Serving-side scaling optimization available.

cache_scale

boolean

Cache scaling optimization available.