Choose a Profile for Your Hardware#
In this tutorial, you select a NIM microservice profile that fits your GPUs using only the NIM Metadata API. The procedure is deterministic, so you can follow it literally, either by hand or in a program.
Select a profile for your GPUs and record the profile_id to supply at
deployment time.
In this tutorial, you will:
Time: Approximately 20 minutes.
Prerequisites#
Before you start, complete the following prerequisites:
Review the Prerequisites and confirm you can reach the base URL.
Complete the Quickstart so that you can fetch a metadata document.
Collect the four hardware inputs described in the following section.
Collect Your Hardware Inputs#
You need four values that describe the hardware you intend to deploy on. The following table describes each value and how to obtain it:
Input |
Description |
How to Obtain It |
|---|---|---|
|
The PCI identifier of your GPU, in the form
|
Run |
|
The vRAM on a single GPU. |
Run |
|
The number of identical GPUs you have available. |
Count the devices you intend to dedicate to this deployment. |
|
The installed NVIDIA driver version. |
Run |
Tutorial Steps#
Follow the steps in order. Each step narrows the candidate set, so the order matters.
Request the metadata document for the NIM microservice and tag you are evaluating, as described in the Quickstart.
If the request returns a 404 status code, first verify the container tag and repository path form (for team-owned repositories, use
nim/{team}/{repo}). If the path form and tag are correct and the response is still 404, this NIM microservice has no published metadata for that tag. Then fall back to the microservice documentation.Success Check: A successful request contains the metadata document for your NIM microservice and tag.
For each entry in
profiles, apply the following tests in order:If
gpu_deviceis present and equals yourgpu_device, record a direct match.If
tested_gpu_devicesis present, is not empty, and contains yourgpu_device, record a tested match.If both
gpu_deviceandgpuare absent, record a GPU-agnostic candidate.Otherwise, discard the profile because it is not a match.
Prefer direct matches over tested matches, and tested matches over GPU-agnostic candidates.
Success Check: Each candidate has a direct, tested, or GPU-agnostic targeting classification.
Discard any profile where
min_vram_per_device_gbis present and greater than yourvram_per_device_gb.If the field is absent, you cannot verify fit from the API. Keep the profile as a candidate, but do not treat it as confirmed.
Success Check: No candidate with a recorded vRAM floor exceeds your per-device vRAM.
Compute the GPU requirement for the profile, treating an absent parallelism value as 1:
required_gpus = (tensor_parallel_size or 1) × (pipeline_parallel_size or 1)
Discard the profile if
required_gpusexceeds yourgpu_count, or ifmax_gpu_countis present and less thanrequired_gpus.Success Check: Every remaining candidate meets the recorded GPU-count constraints.
Evaluate
min_cpu_count,system_memory_per_device_gb,disk_space_per_device_gb, andmin_driver_version.For each field, discard the profile only when the field is present and your hardware falls short.
Success Check: No candidate has a recorded host requirement that your hardware fails.
Filter on the capabilities your workload needs:
If you need LoRA adapters, keep only profiles where
feat_loraistrue.If you need a specific backend, filter on
inference_backend.
An absent flag is unknown rather than false, so decide explicitly whether to keep or discard unknowns.
Success Check: Every remaining candidate meets your capability requirements or has an explicit unknown outcome.
Among the remaining profiles, prefer lower precision for throughput, such as fp8 over bf16 over fp16, or higher precision for accuracy, according to your needs.
Where
throughputis present on all candidates, it can break ties. Review Current Limitations before you rely on it, because its units are not standardized.Success Check: You have ranked the surviving candidates according to your workload needs.
Take
profile_idfrom your chosen profile.That is the value you supply at deployment time.
Success Check: You have recorded the chosen
profile_idfor deployment.
Worked Example#
Suppose you have two L40S GPUs with 48 GB each and the PCI identifier
26b9:10de, and you query a NIM microservice that returns four profiles:
Profile Identifier |
GPU Device |
Minimum vRAM per Device (GB) |
Tensor Parallel Size |
Precision |
|---|---|---|---|---|
|
|
80.0 |
1 |
fp8 |
|
|
48.0 |
2 |
bf16 |
|
|
48.0 |
4 |
fp8 |
|
Absent |
24.0 |
1 |
bf16 |
Working the steps produces the following outcome:
Step 2:
bbbbandccccare direct matches, andddddis a GPU-agnostic candidate. The profileaaaatargets a different device, so it is discarded. Because direct matches exist,dddddrops out of contention.Step 3: Both survivors require 48 GB per device, and you have 48 GB, so both pass.
Step 4: The profile
bbbbrequires two GPUs and you have two, so it passes. The profileccccrequires four, so it is discarded.Steps 5 through 7: The profile
bbbbis the only survivor.
The result is bbbb. Note that cccc was the higher-performance option on
paper, because it uses fp8 rather than bf16, and GPU count alone eliminated it.
This is exactly the kind of determination that previously required downloading
the container to discover.
A Minimal Robust Client#
Install the requests library before you run the following example:
python -m pip install requests
The following example fetches a metadata document with retries and evaluates
hardware fit. Substitute the resourceId and tag for the NIM microservice you
want to inspect:
import requests
from requests.adapters import HTTPAdapter
from urllib3.util.retry import Retry
BASE_URL = "https://api.ngc.nvidia.com/v2"
def _session():
session = requests.Session()
retry = Retry(
total=3,
backoff_factor=1,
status_forcelist=[429, 500, 502, 503, 504],
allowed_methods=["GET"],
)
session.mount("https://", HTTPAdapter(max_retries=retry))
return session
def get_nim_metadata(resource_id, tag):
"""Return the metadata document, or None if none is published."""
response = _session().get(
f"{BASE_URL}/nim/metadata",
params={"resourceId": resource_id, "tag": tag},
headers={"accept": "application/json"},
timeout=30,
)
if response.status_code == 404:
return None
response.raise_for_status()
return response.json()
def fits(profile, vram_per_device_gb, gpu_count):
"""True if the profile fits, False if it does not, None if undeterminable."""
required_vram = profile.get("min_vram_per_device_gb")
if required_vram is None:
return None
if required_vram > vram_per_device_gb:
return False
required_gpus = (
(profile.get("tensor_parallel_size") or 1)
* (profile.get("pipeline_parallel_size") or 1)
)
return required_gpus <= gpu_count
metadata = get_nim_metadata(
"nim/nvidia/parakeet-1-1b-rnnt-multilingual", "1.5.0"
)
if metadata is None:
print("No metadata published")
else:
for profile in metadata.get("profiles", []):
print(
profile.get("profile_id"),
fits(profile, vram_per_device_gb=80, gpu_count=1),
)
The fits function returns three values rather than two. A profile whose vRAM
requirement is unrecorded is genuinely undeterminable, and collapsing that
result into False would silently discard workable profiles.
Troubleshooting#
If a profile matches on gpu but not on gpu_device, match on
gpu_device, because gpu formats vary and are not reliable join keys. If
only gpu is present, you cannot confirm a match programmatically. For other
issues, refer to Troubleshooting.
What You Learned#
In this tutorial, you completed the following tasks:
Collected the four hardware inputs that drive profile selection.
Classified profiles into direct, tested, and GPU-agnostic tiers.
Applied vRAM, GPU-count, host, and functional constraints in order.
Identified the
profile_idto supply at deployment time.
Next Steps#
Continue your learning with these related resources:
Build an integration that tolerates schema changes and partial coverage.
Look up hardware requirements and capability fields.
Understand the constraints that affect profile selection.