NIM Metadata API Documentation#

Every NVIDIA NIM microservice ships with a set of profiles, which are precompiled configurations tuned for particular GPUs, precisions, and inference backends. A single NIM microservice can carry a handful of profiles or several dozen, and the profile you run depends on the hardware you have: how much vRAM per device, how many GPUs, and which driver. The NIM Metadata API answers that question over HTTP, in one request, before you download anything.

Historically, the only complete answer lived in the NIM microservice manifest. That manifest lived inside the container or in the documentation support matrix (for example, the NIM LLM and VLM support matrix). To find out whether a NIM microservice fits your hardware, you had to pull the image, start it, and inspect it. That is a significant investment of time and bandwidth to answer a question you had before you started.

Note

This API is the canonical source for profile metadata. The NIM build process populates it at publish time, so what you read is what the NIM microservice actually contains.

Skip Ahead#

Explore the following information to start working with the NIM Metadata API.

Architecture Overview

Understand publication-time metadata and the HTTP request flow.

Architecture Overview for NIM Metadata API
Ecosystem Overview

Choose between tag-level metadata and complementary NGC information.

NIM Metadata API in the NVIDIA NIM Ecosystem
Get Started

Choose your path through the NIM Metadata API documentation.

About Getting Started with NIM Metadata API
Quickstart

Make your first request and read the response.

Quickstart for NIM Metadata API
Reference Overview

Look up endpoints, response fields, and current limitations.

Reference Overview for NIM Metadata API
Troubleshooting

Resolve unexpected status codes and missing fields.

Troubleshooting NIM Metadata API

Use Cases#

Explore the common use cases for the NIM Metadata API.

Check Fit Before You Pull

Compare the minimum vRAM per device across profiles against the GPUs you have, before you spend bandwidth on a container image.

Choose a Profile for Your Hardware
Choose a Profile Deliberately

Compare precision, inference backend, and tensor or pipeline parallelism side by side instead of accepting a default.

Choose a Profile for Your Hardware
Compare NIM Microservices

Query several NIM microservices and rank them by hardware footprint, context window, or feature support.

Response Fields for NIM Metadata API
Automate Selection

Consume stable, machine-readable JSON from tooling and agents without a human in the loop.

Notes for Automated Agents

Core Concepts#

Explore the core concepts to understand the NIM Metadata API.

Profiles

Deployable configurations that pair a model with a GPU target, a precision, an inference backend, and a parallelism layout.

Core Concepts Overview for NIM Metadata API
GPU Targeting

How the API distinguishes a profile that targets a specific GPU from one that runs anywhere with sufficient vRAM.

Core Concepts Overview for NIM Metadata API
Absent Means Unknown

Why the service omits empty fields entirely, and why an absent field is never the same as a false or zero value.

Core Concepts Overview for NIM Metadata API
Metadata Coverage

Why a 404 response is a normal outcome rather than a fault in your request.

NIM Metadata API Prerequisites

Scope of This Documentation#

This documentation covers the public read endpoints, which use the HTTP GET method. Reading metadata is not the same as gaining access to a NIM microservice. Being able to read the profiles for a NIM microservice does not grant you the right to pull the container. Pulling still follows the normal entitlement rules for that container.