Support Matrix for NIMs#
This page lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM.
Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.
NVIDIA publishes inference microservices (NIMs) under distinct NIM offerings. Models on this page are grouped in separate sections based on NIM offering: NIM Certified and NIM. For more information, refer to NIM Offerings.
Use the table below to filter NIM profiles by GPU, tensor parallelism (TP), precision, certification status, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.
NIM Certified#
NIM Certified is the enterprise production NIM offering. It requires NVIDIA AI Enterprise.
gpt-oss-120b#
The following table lists the supported profile configurations for openai/gpt-oss-120b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
MXFP4 |
|
|
|
|
MXFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
gpt-oss-20b#
The following table lists the supported profile configurations for openai/gpt-oss-20b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
MXFP4 |
|
|
|
|
MXFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.1-70b-instruct#
The following table lists the supported profile configurations for meta/llama-3.1-70b-instruct:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A10GNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.1-8b-instruct#
The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct:
Precision |
TP1 |
|---|---|
BF16 |
|
BF16 + LoRA |
|
FP8 |
|
FP8 + LoRA |
|
NVFP4 |
|
NVFP4 + LoRA |
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.3-70b-instruct#
The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A10GNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.3-nemotron-super-49b-v1.5#
The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A10GNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
nemotron-3-nano#
The following table lists the supported profile configurations for nvidia/nemotron-3-nano:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
– |
– |
– |
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
nemotron-3-super-120b-a12b#
The following table lists the supported profile configurations for nvidia/nemotron-3-super-120b-a12b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
– |
|
|
|
FP8 |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
– |
– |
Note
This is a large model. Lower-TP profiles require substantially more GPU memory per device, so some verified GPUs support only TP4 or TP8 profiles.
Verified GPUs
The following GPUs have been verified with one or more supported profiles for this model:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
nemotron-3-ultra-550b-a55b#
The following table lists the supported profile configurations for nvidia/nemotron-3-ultra-550b-a55b:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
– |
– |
– |
|
BF16 + LoRA |
– |
– |
– |
|
NVFP4 |
– |
|
|
|
Note
The minimum deployment is two GPUs, using vllm-nvfp4-tp2-pp1 on high-memory Blackwell SKUs (NVIDIA-B300-SXM6-AC and NVIDIA-GB300); the other profiles use four or eight GPUs. H200 validation covers one NVFP4 TP4 profile and is less extensive than the Blackwell validation.
Verified GPUs
The following GPUs have been verified with one or more supported profiles for this model:
NVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-GB300NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200
riva-translate-4b-instruct-v2#
The following table lists the supported profile configuration for nvidia/riva-translate-4b-instruct-v2:
Precision |
TP |
PP |
LoRA |
Profile |
|---|---|---|---|---|
BF16 |
1 |
1 |
No |
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
starcoder2-7b#
The following table lists the supported profile configurations for bigcode/starcoder2-7b:
Precision |
TP1 |
TP2 |
|---|---|---|
BF16 |
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-GB300-WSNVIDIA-H100-80GB-HBM3NVIDIA-H200
Model-Free NIM#
The following models are tested and validated for
nvidia/model-free-nim:
gpt-oss-20bapriel-nemotroncodestral
While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.
Verified GPUs
The model-free NIM has been verified on the following GPUs:
NVIDIA-A100-80GB-PCIeNVIDIA-A100-PCIE-40GBNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B200NVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-GB300NVIDIA-GB300-WSNVIDIA-GH200-144G-HBM3eNVIDIA-GH200-480GBNVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H100-PCIeNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
NIM#
NIM is the offering that delivers NIMs that are validated to be functional on a small set of NVIDIA GPUs. It is free to use and is not part of the NVIDIA AI Enterprise portfolio.
nemotron-content-safety-reasoning-4b-experimental#
The following table lists the supported profile configurations for
nvidia/nemotron-content-safety-reasoning-4b-experimental:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-B200NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H200
For information on getting started with nemotron-content-safety-reasoning-4b-experimental, refer to version 2.0.3 of the documentation.
1.x NIM LLM Models#
For more information on version 1.x NIMs, refer to the 1.15 version of the NIM LLM Supported Models page.
Show 1.x models
Model (Hardware Requirements) |
Organization/Model ID (Catalog Page) |
|---|---|
|
|