NIM Metadata API Documentation#
Every NVIDIA NIM microservice ships with a set of profiles, which are precompiled configurations tuned for particular GPUs, precisions, and inference backends. A single NIM microservice can carry a handful of profiles or several dozen, and the profile you run depends on the hardware you have: how much vRAM per device, how many GPUs, and which driver. The NIM Metadata API answers that question over HTTP, in one request, before you download anything.
Historically, the only complete answer lived in the NIM microservice manifest. That manifest lived inside the container or in the documentation support matrix (for example, the NIM LLM and VLM support matrix). To find out whether a NIM microservice fits your hardware, you had to pull the image, start it, and inspect it. That is a significant investment of time and bandwidth to answer a question you had before you started.
Note
This API is the canonical source for profile metadata. The NIM build process populates it at publish time, so what you read is what the NIM microservice actually contains.
Skip Ahead#
Explore the following information to start working with the NIM Metadata API.
Understand publication-time metadata and the HTTP request flow.
Choose between tag-level metadata and complementary NGC information.
Choose your path through the NIM Metadata API documentation.
Make your first request and read the response.
Look up endpoints, response fields, and current limitations.
Resolve unexpected status codes and missing fields.
Use Cases#
Explore the common use cases for the NIM Metadata API.
Compare the minimum vRAM per device across profiles against the GPUs you have, before you spend bandwidth on a container image.
Compare precision, inference backend, and tensor or pipeline parallelism side by side instead of accepting a default.
Query several NIM microservices and rank them by hardware footprint, context window, or feature support.
Consume stable, machine-readable JSON from tooling and agents without a human in the loop.
Core Concepts#
Explore the core concepts to understand the NIM Metadata API.
Deployable configurations that pair a model with a GPU target, a precision, an inference backend, and a parallelism layout.
How the API distinguishes a profile that targets a specific GPU from one that runs anywhere with sufficient vRAM.
Why the service omits empty fields entirely, and why an absent field is never the same as a false or zero value.
Why a 404 response is a normal outcome rather than a fault in your request.
Scope of This Documentation#
This documentation covers the public read endpoints, which use the HTTP GET
method. Reading metadata is not the same as gaining access to a NIM
microservice. Being able to read the profiles for a NIM microservice does not
grant you the right to pull the container. Pulling still follows the normal
entitlement rules for that container.