Support Matrix for PB NIMs#
This page lists the supported models, their deployment profiles, and the verified hardware SKUs for NIM LLM NIMs on the Production Branch (PB).
Supported Models and Profiles#
Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.
Note
Some GPU SKUs can encounter out-of-memory (OOM) errors on memory-constrained hardware. If a model fails to start due to insufficient GPU memory, see Troubleshooting GPU Memory Out-of-Memory Errors for mitigations such as reducing --max-model-len.
Use the table below to filter certified NIM profiles by GPU, tensor parallelism (TP), precision, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.
llama-3.1-8b-instruct-pb6#
Latest supported NIM LLM PB version: 2.0.4-pb6.3
The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct-pb6:
Precision |
TP1 |
|---|---|
BF16 |
|
BF16 + LoRA |
|
FP8 |
|
FP8 + LoRA |
|
NVFP4 |
|
NVFP4 + LoRA |
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.3-70b-instruct-pb6#
Latest supported NIM LLM PB version: 2.0.4-pb6.3
The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct-pb6:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A10GNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
llama-3.3-nemotron-super-49b-v1.5-pb6#
Latest supported NIM LLM PB version: 2.0.4-pb6.3
The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5-pb6:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
NVFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A10GNVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
nemotron-3-nano-pb6#
Latest supported NIM LLM PB version: 2.0.4-pb6.3
The following table lists the supported profile configurations for nvidia/nemotron-3-nano-pb6:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
BF16 |
|
|
|
|
BF16 + LoRA |
|
|
|
|
FP8 |
|
|
|
|
FP8 + LoRA |
|
|
|
|
NVFP4 |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
gpt-oss-120b-pb6#
Latest supported NIM LLM PB version: 2.0.4-pb6.3
The following table lists the supported profile configurations for openai/gpt-oss-120b-pb6:
Precision |
TP1 |
TP2 |
TP4 |
TP8 |
|---|---|---|---|---|
MXFP4 |
|
|
|
|
MXFP4 + LoRA |
|
|
|
|
Verified GPUs
This model has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B300-SXM6-ACNVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition
model-free-nim-pb6#
Latest supported NIM LLM PB version: 2.0.4-pb6.3
The following models are tested and validated for
nvidia/model-free-nim-pb6:
gpt-oss-20bapriel-nemotroncodestral
While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.
Verified GPUs
The model-free NIM has been verified on the following GPUs:
NVIDIA-A100-SXM4-40GBNVIDIA-A100-SXM4-80GBNVIDIA-A10GNVIDIA-B300-SXM6-ACNVIDIA-GB10NVIDIA-GB200NVIDIA-H100-80GB-HBM3NVIDIA-H100-NVLNVIDIA-H200NVIDIA-H200-NVLNVIDIA-L40SNVIDIA-RTX-PRO-4500-Blackwell-Server-EditionNVIDIA-RTX-PRO-6000-Blackwell-Server-Edition