Quickstart for NIM Metadata API#

This Quickstart walks you through identifying a NIM microservice, making your first request, and reading the response. For the requirements you need first, refer to Prerequisites.

Prerequisites#

Before you start, complete the following prerequisites:

  1. Install an HTTP client, such as curl or the Python requests library.

  2. Collect the resourceId and tag for the NIM microservice you want to inspect, as described in Prerequisites.

Quickstart Steps#

To quickly get started with the NIM Metadata API, complete the following steps.

  1. Identify the NIM microservice by its resourceId and tag.

    The resourceId is a slash-separated repository path, and the tag is the container tag as it appears in the catalog. For this example, use nim/nvidia/parakeet-1-1b-rnnt-multilingual and 1.5.0.

  2. Make the request against the query-parameter form of the endpoint.

    curl --request GET \
      --url 'https://api.ngc.nvidia.com/v2/nim/metadata?resourceId=nim/nvidia/parakeet-1-1b-rnnt-multilingual&tag=1.5.0' \
      --header 'accept: application/json'
    

    Install the requests library before you run the following example:

    python -m pip install requests
    

    The following example makes the request and reports the profile count:

    import requests
    
    BASE_URL = "https://api.ngc.nvidia.com/v2"
    
    response = requests.get(
        f"{BASE_URL}/nim/metadata",
        params={"resourceId": "nim/nvidia/parakeet-1-1b-rnnt-multilingual", "tag": "1.5.0"},
        headers={"accept": "application/json"},
        timeout=30,
    )
    response.raise_for_status()
    metadata = response.json()
    print(f"{len(metadata.get('profiles', []))} profiles")
    

    A 200 status code returns the metadata document. A 404 status code means this combination of NIM microservice and tag has no published metadata.

  3. Read the four top-level sections of the response.

    All four are optional.

    • profiles: The deployable configurations. This is the section you use most.

    • release_details: Release-level facts, such as version, documentation, compliance, and supported architectures.

    • features: Capability flags, such as context window, tool calling, and backend support.

    • related_urls: Links to related resources.

  4. Compare the profiles against your hardware.

    The following abridged response shows a NIM microservice with GPU-specific profiles:

    {
      "profiles": [
        {
          "profile_id": "a1b2c3d4e5f60718293a4b5c6d7e8f90",
          "gpu": "H100",
          "gpu_device": "2330:10de",
          "tested_gpu_devices": ["2330:10de"],
          "precision": "fp8",
          "inference_backend": "tensorrt_llm",
          "tensor_parallel_size": 1,
          "min_vram_per_device_gb": 80.0,
          "feat_lora": false
        },
        {
          "profile_id": "0f1e2d3c4b5a69788796a5b4c3d2e1f0",
          "gpu": "L40S",
          "gpu_device": "26b9:10de",
          "tested_gpu_devices": ["26b9:10de"],
          "precision": "bf16",
          "inference_backend": "vllm",
          "tensor_parallel_size": 2,
          "min_vram_per_device_gb": 48.0,
          "feat_lora": false
        }
      ],
      "release_details": {
        "container_version": "1.2.0",
        "architecture": ["AMD 64", "ARM 64"],
        "documentation_url": "https://docs.nvidia.com/nim/"
      },
      "features": {
        "context_window": 8192,
        "tool_calling": false,
        "vllm_support": true
      }
    }
    

    This response describes two profiles, two different GPUs, and two different vRAM floors. If you have L40S GPUs with 48 GB each, the second profile fits and the first does not. That comparison, made in one request before you download anything, is the core purpose of this API.

  5. Recognize the GPU-agnostic case.

    Many NIM microservices ship profiles that run anywhere with sufficient vRAM. Those profiles omit gpu and gpu_device entirely, and can carry an empty tested_gpu_devices array:

    {
      "profiles": [
        {
          "profile_id": "7c8d9e0f1a2b3c4d5e6f708192a3b4c5",
          "tested_gpu_devices": [],
          "precision": "bf16",
          "inference_backend": "vllm",
          "tensor_parallel_size": 1
        }
      ]
    }
    

    This is not an empty or broken response. Read it as “no GPUs tested with the profile” rather than “no GPUs supported.” For the full set of GPU targeting signals, refer to Core Concepts Overview.

Minimal Code Example#

Use the following self-contained example to fetch a metadata document and report the profiles that fit a given vRAM budget:

import requests

BASE_URL = "https://api.ngc.nvidia.com/v2"


def get_nim_metadata(resource_id, tag):
    """Return the metadata document, or None if none is published."""
    response = requests.get(
        f"{BASE_URL}/nim/metadata",
        params={"resourceId": resource_id, "tag": tag},
        headers={"accept": "application/json"},
        timeout=30,
    )
    if response.status_code == 404:
        return None
    response.raise_for_status()
    return response.json()


metadata = get_nim_metadata("nim/nvidia/parakeet-1-1b-rnnt-multilingual", "1.5.0")
if metadata is None:
    print("No published metadata for this NIM microservice and tag.")
else:
    for profile in metadata.get("profiles", []):
        required_vram = profile.get("min_vram_per_device_gb")
        print(profile["profile_id"], required_vram)

Next Steps#

Now that you have a working request, explore the following resources in greater detail: