Notes for Automated Agents#

This page is written for programmatic consumers of the NIM Metadata API, such as tooling, schedulers, and agents that select profiles without a human in the loop. If you are exploring the API by hand, start with the Quickstart instead.

The response is stable, machine-readable JSON, but the schema grows over time. Following the rules on this page keeps an integration working across schema changes rather than breaking on the first new field.

Forward-Compatibility Rules#

Apply the following six rules to every integration.

Rule

Description

  1. Never assume a field is present

Only profile_id and tested_gpu_devices are guaranteed present. Access everything else defensively, with an explicit branch for absence.

  1. Never treat absence as a value

Absent is not false, 0, or an empty string. If a decision depends on an absent field, surface the outcome as unknown rather than defaulting it.

  1. Ignore unrecognized fields

New fields are added over time. Deserialize permissively and never fail on an unknown key.

  1. Tolerate gpu and gpu_device becoming arrays

A single profile_id can be valid on multiple compatible GPUs. These fields can change from a string to a list, so accept either shape.

  1. Treat precision and inference_backend as open strings

These fields are not enumerations. Carry through new values instead of failing on an unrecognized value.

  1. Do not parse statusDescription

Branch on the HTTP status code and requestStatus.statusCode. The description text is written for people and changes over time.

The following example accepts either GPU field shape. Complete the Quickstart first to define metadata:

def as_list(value):
    if value is None:
        return []
    return value if isinstance(value, list) else [value]


for profile in metadata.get("profiles", []):
    gpus = as_list(profile.get("gpu"))
    devices = as_list(profile.get("gpu_device"))
    print(profile.get("profile_id"), gpus, devices)

Freshness and Caching#

The API offers no ETag header, no Last-Modified header, and no conditional-request support, so cache based on tag stability rather than cache validation. The following table describes how long to trust a cached document:

Tag Type

Stability

Guidance

Semantic tags, such as 1.2.0

Stable in practice, but not guaranteed immutable

Cache for hours or days, not forever. The service refuses to create metadata over a tag that already has it, but the publishing pipeline can still amend or remove it.

Rolling tags, such as latest

Not stable

Do not cache beyond your tolerance for staleness, because these tags are republished.

Publication lags the container. A newly pushed container tag can exist before its metadata does, so a 404 status code immediately after a release can resolve shortly afterward. If you poll, use a bounded number of attempts with backoff rather than polling indefinitely.

Response Sizes#

Plan for variability rather than a fixed size. A typical NIM microservice returns roughly 8 KB across eight profiles. NIM microservices with many profiles reach 20 KB to 50 KB, and the largest observed historical documents approach 300 KB across several hundred profiles.

There is no hard cap on profile count, so stream the response or size your buffers accordingly and do not assume a small fixed maximum.

Rate Limits#

No endpoint-specific rate limit is published, and standard NGC gateway limits apply. Be a considerate client by caching immutable tags, avoiding tight polling loops, and backing off on a 429 status code or any 5xx status code.

Best Practices#

Apply the following practices when you build against this API:

  • Branch on status codes, not text. Treat a 404 status code as a definitive answer about metadata availability rather than an error to retry.

  • Model unknowns explicitly. Return a three-valued result from fit checks so that an unrecorded requirement is distinguishable from a failed one.

  • Log the request identifier. Record requestStatus.requestId on every non-success response, because that is the value support needs to trace a specific request.

  • Do not retry service-side data problems. A well-formed request that returns a 400 or 500 status code because of stored data will not succeed on retry. Report it instead.

  • Design for partial coverage. Metadata is attached at publish time, so build the no-metadata path first rather than treating it as an edge case.

Next Steps#

Explore the following resources to complete your integration: