Support Matrix for NIMs#
This page lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM.
Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.
Important
The list of NIM Certified models is available on the NGC Catalog site.
Use the table below to filter NIM profiles by GPU, tensor parallelism (TP), precision, LoRA support, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.
Note
The GB300-WS GPU is more commonly known as DGX Station.
gpt-oss-120b#
The following table lists the supported profile configurations for openai/gpt-oss-120b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
MXFP4 |
|
|
|
|
MXFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
Note
Support requirements in 2.0.11:
The MXFP4 TP8 profile on NVIDIA A10G and NVIDIA A100-SXM4-40GB requires
NIM_KVCACHE_PERCENT=0.80. The LoRA TP4 and TP8 profiles on NVIDIA L40S and
NVIDIA RTX PRO 4500 Blackwell Server Edition require piecewise CUDA graphs.
LoRA adapters with rank greater than 128 are not supported.
gpt-oss-20b#
The following table lists the supported profile configurations for openai/gpt-oss-20b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
MXFP4 |
|
|
|
|
MXFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
Note
Support requirements in 2.0.11:
LoRA adapters with rank greater than 128 are not supported. On NVIDIA A10G,
LoRA deployments require NIM_MAX_MODEL_LEN below the default 131,072 token
context length so that the KV cache fits in device memory.
llama-3.1-70b-instruct#
The following table lists the supported profile configurations for meta/llama-3.1-70b-instruct:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
Note
Support change in 2.0.11: BF16 TP2 with or without LoRA is not supported on NVIDIA RTX PRO 6000 Blackwell Server Edition, and BF16 TP8 without LoRA is not supported on NVIDIA B200. These combinations are omitted from the profile-to-GPU map above.
llama-3.1-8b-instruct#
The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct:
Precision |
TP1 |
|---|---|
BF16 |
|
BF16 + LoRA |
|
FP8 |
|
FP8 + LoRA |
|
NVFP4 |
|
NVFP4 + LoRA |
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.3-70b-instruct#
The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.3-nemotron-super-49b-v1.5#
The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
nemotron-content-safety-reasoning-4b-experimental#
The following table lists the supported profile configurations for
nvidia/nemotron-content-safety-reasoning-4b-experimental. The latest
release version of this NIM is 2.0.3.
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-B200NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H200
nemotron-3-nano#
The following table lists the supported profile configurations for nvidia/nemotron-3-nano:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
– |
– |
– |
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
Note
Support requirement in 2.0.11: The fused mixture-of-experts kernels support LoRA adapters with a maximum rank of 128. Adapters with a greater rank are not supported.
nemotron-3-super-120b-a12b#
The following table lists the supported profile configurations for nvidia/nemotron-3-super-120b-a12b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
– |
|
|
|
FP8 |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
– |
– |
Note
This is a large model. Lower-TP profiles require substantially more GPU memory per device, so some verified GPUs support only TP4 or TP8 profiles.
Verified GPUs
The following GPUs have been verified with one or more supported profiles for this model:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
Note
Support change in 2.0.11: The following BF16 LoRA combinations are not supported: TP2 on NVIDIA GH200-144G-HBM3e and NVIDIA H200-NVL, TP4 on NVIDIA H100-NVL, and TP8 on NVIDIA L40S. These combinations are omitted from the profile-to-GPU map above.
nemotron-3-ultra-550b-a55b#
The following table lists the supported profile configurations for nvidia/nemotron-3-ultra-550b-a55b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
– |
– |
– |
|
BF16 + LoRA |
– |
– |
– |
– |
NVFP4 |
– |
|
|
|
Note
The minimum deployment is two GPUs, using vllm-nvfp4-tp2-pp1 on high-memory Blackwell SKUs (NVIDIA-B300-SXM6-AC and NVIDIA-GB300); the other profiles use four or eight GPUs. H200 validation covers one NVFP4 TP4 profile and is less extensive than the Blackwell validation.
FP8 profiles were not released for this NIM because NVFP4 profiles work on all tested Hopper-architecture GPUs (for example, H100 and H200) and deliver better performance than FP8 profiles.
Verified GPUs
The following GPUs have been verified with one or more supported profiles for this model:
NVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200
nemotron-3.5-lightning-30b-a3b#
The following table lists the supported profile configurations for
nvidia/nemotron-3.5-lightning-30b-a3b, along
with the minimum GPU VRAM required per device and the GPU architecture that
each precision requires. The VRAM values come from the profile disk-sizing
metadata reported by list-model-profiles and represent the minimum per-GPU
memory needed to load the model weights and the runtime allocations; they do
not include additional headroom for large KV caches at high context length or
high concurrency.
The latest release version of this NIM is 2.0.10.
Precision |
TP |
PP |
Min VRAM per GPU |
Min GPU count |
Architecture requirement |
LoRA Profile |
|---|---|---|---|---|---|---|
BF16 |
1 |
1 |
66 GB |
1 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
BF16 |
2 |
1 |
35 GB |
2 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
BF16 |
4 |
1 |
20 GB |
4 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
BF16 |
8 |
1 |
12 GB |
8 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
W4A16 |
1 |
1 |
32 GB |
1 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
W4A16 |
2 |
1 |
20 GB |
2 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
W4A16 |
4 |
1 |
14 GB |
4 |
Ampere or newer (SM 8.0+) |
Base and LoRA |
NVFP4 |
1 |
1 |
30 GB |
1 |
Blackwell or newer (SM 10.0+) |
Base and LoRA |
NVFP4 |
2 |
1 |
18 GB |
2 |
Blackwell or newer (SM 10.0+) |
Base and LoRA |
NVFP4 |
4 |
1 |
12 GB |
4 |
Blackwell or newer (SM 10.0+) |
Base and LoRA |
Notes:
Min VRAM per GPU is the floor that
list-model-profilesreports for that precision and tensor-parallel size. A GPU below the floor is filtered out during profile selection.Architecture requirement reflects the kernels each precision needs. NVFP4 requires Streaming Multiprocessor (SM) 10.0+ (Blackwell) because it uses NVFP4-native tensor cores. BF16 and W4A16 run on any Ampere-class or newer GPU (SM 8.0+).
Sharing weights across more GPUs (higher TP) lowers the per-GPU VRAM floor and enables lower-memory GPUs.
Model-cache disk sizes range from approximately 19 GB (NVFP4) to approximately 63 GB (BF16). Cache size is separate from GPU VRAM and does not indicate that the model will fit in a given GPU.
Verified GPUs
This model has been verified on the following GPUs (grouped by architecture):
Blackwell (SM 10.0) — supports NVFP4, W4A16, BF16:
GPU |
VRAM per GPU |
Profiles that fit |
|---|---|---|
|
180 GB |
All profiles |
|
288 GB |
All profiles |
|
186 GB |
All profiles |
|
288 GB |
All profiles |
|
128 GB (unified memory) |
All profiles; use |
|
96 GB |
All profiles |
Hopper (SM 9.0) — supports W4A16 and BF16 (no NVFP4):
GPU |
VRAM per GPU |
Profiles that fit |
|---|---|---|
|
80 GB |
BF16 all sizes (needs |
|
141 GB |
BF16 all sizes, W4A16 all sizes |
Ampere (SM 8.0) — supports W4A16 and BF16 (no NVFP4):
GPU |
VRAM per GPU |
Profiles that fit |
|---|---|---|
|
80 GB |
BF16 TP≥2, W4A16 all sizes |
Ada (SM 8.9) — supports W4A16 and BF16 (no NVFP4):
GPU |
VRAM per GPU |
Profiles that fit |
|---|---|---|
|
48 GB |
BF16 TP≥2, W4A16 all sizes |
Unverified but potentially compatible. Any GPU that meets the per-profile minimum VRAM floor and the architecture requirement in the profile table can run that profile technically, even if it is not in the preceding Verified GPUs tables. Only the GPUs in those tables have been tested end-to-end for this model at this NIM version; untested systems may exhibit different startup, throughput, or accuracy behavior.
riva-translate-4b-instruct-v2#
The following table lists the supported profile configuration for nvidia/riva-translate-4b-instruct-v2.
The latest release version of this NIM is
2.0.8.
Precision |
TP |
PP |
LoRA |
Profile |
|---|---|---|---|---|
BF16 |
1 |
1 |
No |
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
starcoder2-7b#
The following table lists the supported profile configurations for bigcode/starcoder2-7b:
Precision |
TP1 |
TP2 |
|---|---|---|
BF16 |
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-80GB-PCIeNVIDIA-A100-SXM4-80GBNVIDIA-GB300-WSNVIDIA-H100-80GB-HBM3NVIDIA-H200
Model-Free NIM#
The following models are tested and validated for
nvidia/model-free-nim:
gpt-oss-20bapriel-nemotroncodestral
While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.
Verified GPUs
The model-free NIM has been verified on the following GPUs:
NVIDIA-A100-80GB-PCIeNVIDIA-A100-PCIE-40GBNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H100-PCIeNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
Note
Model-specific VRAM requirement in 2.0.11:
When this NIM serves apriel-nemotron or codestral, use a GPU with at least
80 GB of usable VRAM. These models require approximately 66 GB for weights and
runtime allocations and cannot start on smaller GPUs. Other validated models
can use the remaining verified GPUs listed above.
1.x NIM LLM Models#
For more information on version 1.x NIMs, refer to the 1.15 version of the NIM LLM Supported Models page.
Show 1.x models
Model (Hardware Requirements) |
Organization/Model ID (Catalog Page) |
|---|---|
|
|