Choose a Profile for Your Hardware#

In this tutorial, you select a NIM microservice profile that fits your GPUs using only the NIM Metadata API. The procedure is deterministic, so you can follow it literally, either by hand or in a program.

Tutorial Goal

Select a profile for your GPUs and record the profile_id to supply at deployment time.

In this tutorial, you will:

  1. Fetch metadata and classify GPU targeting.

  2. Apply hardware constraints.

  3. Filter capabilities and rank candidates.

  4. Record the selected profile identifier.

Time: Approximately 20 minutes.

Prerequisites#

Before you start, complete the following prerequisites:

  1. Review the Prerequisites and confirm you can reach the base URL.

  2. Complete the Quickstart so that you can fetch a metadata document.

  3. Collect the four hardware inputs described in the following section.

Collect Your Hardware Inputs#

You need four values that describe the hardware you intend to deploy on. The following table describes each value and how to obtain it:

Input

Description

How to Obtain It

gpu_device

The PCI identifier of your GPU, in the form pciDeviceId:pciVendorId.

Run lspci -nn | grep -i nvidia on Linux. The output shows [10de:2330], which reverses to 2330:10de.

vram_per_device_gb

The vRAM on a single GPU.

Run nvidia-smi --query-gpu=memory.total --format=csv.

gpu_count

The number of identical GPUs you have available.

Count the devices you intend to dedicate to this deployment.

driver_version

The installed NVIDIA driver version.

Run nvidia-smi --query-gpu=driver_version --format=csv.

Tutorial Steps#

Follow the steps in order. Each step narrows the candidate set, so the order matters.

  1. Request the metadata document for the NIM microservice and tag you are evaluating, as described in the Quickstart.

    If the request returns a 404 status code, first verify the container tag and repository path form (for team-owned repositories, use nim/{team}/{repo}). If the path form and tag are correct and the response is still 404, this NIM microservice has no published metadata for that tag. Then fall back to the microservice documentation.

    Success Check: A successful request contains the metadata document for your NIM microservice and tag.

  2. For each entry in profiles, apply the following tests in order:

    • If gpu_device is present and equals your gpu_device, record a direct match.

    • If tested_gpu_devices is present, is not empty, and contains your gpu_device, record a tested match.

    • If both gpu_device and gpu are absent, record a GPU-agnostic candidate.

    • Otherwise, discard the profile because it is not a match.

    Prefer direct matches over tested matches, and tested matches over GPU-agnostic candidates.

    Success Check: Each candidate has a direct, tested, or GPU-agnostic targeting classification.

  3. Discard any profile where min_vram_per_device_gb is present and greater than your vram_per_device_gb.

    If the field is absent, you cannot verify fit from the API. Keep the profile as a candidate, but do not treat it as confirmed.

    Success Check: No candidate with a recorded vRAM floor exceeds your per-device vRAM.

  4. Compute the GPU requirement for the profile, treating an absent parallelism value as 1:

    required_gpus = (tensor_parallel_size or 1) × (pipeline_parallel_size or 1)
    

    Discard the profile if required_gpus exceeds your gpu_count, or if max_gpu_count is present and less than required_gpus.

    Success Check: Every remaining candidate meets the recorded GPU-count constraints.

  5. Evaluate min_cpu_count, system_memory_per_device_gb, disk_space_per_device_gb, and min_driver_version.

    For each field, discard the profile only when the field is present and your hardware falls short.

    Success Check: No candidate has a recorded host requirement that your hardware fails.

  6. Filter on the capabilities your workload needs:

    • If you need LoRA adapters, keep only profiles where feat_lora is true.

    • If you need a specific backend, filter on inference_backend.

    An absent flag is unknown rather than false, so decide explicitly whether to keep or discard unknowns.

    Success Check: Every remaining candidate meets your capability requirements or has an explicit unknown outcome.

  7. Among the remaining profiles, prefer lower precision for throughput, such as fp8 over bf16 over fp16, or higher precision for accuracy, according to your needs.

    Where throughput is present on all candidates, it can break ties. Review Current Limitations before you rely on it, because its units are not standardized.

    Success Check: You have ranked the surviving candidates according to your workload needs.

  8. Take profile_id from your chosen profile.

    That is the value you supply at deployment time.

    Success Check: You have recorded the chosen profile_id for deployment.

Worked Example#

Suppose you have two L40S GPUs with 48 GB each and the PCI identifier 26b9:10de, and you query a NIM microservice that returns four profiles:

Profile Identifier

GPU Device

Minimum vRAM per Device (GB)

Tensor Parallel Size

Precision

aaaa…

2330:10de

80.0

1

fp8

bbbb…

26b9:10de

48.0

2

bf16

cccc…

26b9:10de

48.0

4

fp8

dddd…

Absent

24.0

1

bf16

Working the steps produces the following outcome:

  • Step 2: bbbb and cccc are direct matches, and dddd is a GPU-agnostic candidate. The profile aaaa targets a different device, so it is discarded. Because direct matches exist, dddd drops out of contention.

  • Step 3: Both survivors require 48 GB per device, and you have 48 GB, so both pass.

  • Step 4: The profile bbbb requires two GPUs and you have two, so it passes. The profile cccc requires four, so it is discarded.

  • Steps 5 through 7: The profile bbbb is the only survivor.

The result is bbbb. Note that cccc was the higher-performance option on paper, because it uses fp8 rather than bf16, and GPU count alone eliminated it. This is exactly the kind of determination that previously required downloading the container to discover.

A Minimal Robust Client#

Install the requests library before you run the following example:

python -m pip install requests

The following example fetches a metadata document with retries and evaluates hardware fit. Substitute the resourceId and tag for the NIM microservice you want to inspect:

import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry

BASE_URL = "https://api.ngc.nvidia.com/v2"


def _session():
    session = requests.Session()
    retry = Retry(
        total=3,
        backoff_factor=1,
        status_forcelist=[429, 500, 502, 503, 504],
        allowed_methods=["GET"],
    )
    session.mount("https://", HTTPAdapter(max_retries=retry))
    return session


def get_nim_metadata(resource_id, tag):
    """Return the metadata document, or None if none is published."""
    response = _session().get(
        f"{BASE_URL}/nim/metadata",
        params={"resourceId": resource_id, "tag": tag},
        headers={"accept": "application/json"},
        timeout=30,
    )
    if response.status_code == 404:
        return None
    response.raise_for_status()
    return response.json()


def fits(profile, vram_per_device_gb, gpu_count):
    """True if the profile fits, False if it does not, None if undeterminable."""
    required_vram = profile.get("min_vram_per_device_gb")
    if required_vram is None:
        return None
    if required_vram > vram_per_device_gb:
        return False
    required_gpus = (
        (profile.get("tensor_parallel_size") or 1)
        * (profile.get("pipeline_parallel_size") or 1)
    )
    return required_gpus <= gpu_count


metadata = get_nim_metadata(
    "nim/nvidia/parakeet-1-1b-rnnt-multilingual", "1.5.0"
)
if metadata is None:
    print("No metadata published")
else:
    for profile in metadata.get("profiles", []):
        print(
            profile.get("profile_id"),
            fits(profile, vram_per_device_gb=80, gpu_count=1),
        )

The fits function returns three values rather than two. A profile whose vRAM requirement is unrecorded is genuinely undeterminable, and collapsing that result into False would silently discard workable profiles.

Troubleshooting#

If a profile matches on gpu but not on gpu_device, match on gpu_device, because gpu formats vary and are not reliable join keys. If only gpu is present, you cannot confirm a match programmatically. For other issues, refer to Troubleshooting.

What You Learned#

In this tutorial, you completed the following tasks:

  • Collected the four hardware inputs that drive profile selection.

  • Classified profiles into direct, tested, and GPU-agnostic tiers.

  • Applied vRAM, GPU-count, host, and functional constraints in order.

  • Identified the profile_id to supply at deployment time.

Next Steps#

Continue your learning with these related resources:

Notes for Automated Agents

Build an integration that tolerates schema changes and partial coverage.

Notes for Automated Agents
Response Fields

Look up hardware requirements and capability fields.

Response Fields for NIM Metadata API
Current Limitations

Understand the constraints that affect profile selection.

Current Limitations of the NIM Metadata API