Support Matrix for PB NIMs#

This page lists the supported models, their deployment profiles, and the verified hardware SKUs for NIM LLM NIMs on the Production Branch (PB).

Supported Models and Profiles#

Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.

Note

Some GPU SKUs can encounter out-of-memory (OOM) errors on memory-constrained hardware. If a model fails to start due to insufficient GPU memory, see Troubleshooting GPU Memory Out-of-Memory Errors for mitigations such as reducing --max-model-len.

Use the table below to filter certified NIM profiles by GPU, tensor parallelism (TP), precision, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.

Model TP Precision LoRA
llama-3.1-8b-instruct-pb6 1 BF16 No
llama-3.1-8b-instruct-pb6 1 BF16 Yes
llama-3.1-8b-instruct-pb6 1 FP8 No
llama-3.1-8b-instruct-pb6 1 FP8 Yes
llama-3.1-8b-instruct-pb6 1 NVFP4 No
llama-3.1-8b-instruct-pb6 1 NVFP4 Yes
llama-3.3-70b-instruct-pb6 1 BF16 No
llama-3.3-70b-instruct-pb6 2 BF16 No
llama-3.3-70b-instruct-pb6 4 BF16 No
llama-3.3-70b-instruct-pb6 8 BF16 No
llama-3.3-70b-instruct-pb6 1 BF16 Yes
llama-3.3-70b-instruct-pb6 2 BF16 Yes
llama-3.3-70b-instruct-pb6 4 BF16 Yes
llama-3.3-70b-instruct-pb6 8 BF16 Yes
llama-3.3-70b-instruct-pb6 1 FP8 No
llama-3.3-70b-instruct-pb6 2 FP8 No
llama-3.3-70b-instruct-pb6 4 FP8 No
llama-3.3-70b-instruct-pb6 8 FP8 No
llama-3.3-70b-instruct-pb6 1 FP8 Yes
llama-3.3-70b-instruct-pb6 2 FP8 Yes
llama-3.3-70b-instruct-pb6 4 FP8 Yes
llama-3.3-70b-instruct-pb6 8 FP8 Yes
llama-3.3-70b-instruct-pb6 1 NVFP4 No
llama-3.3-70b-instruct-pb6 2 NVFP4 No
llama-3.3-70b-instruct-pb6 4 NVFP4 No
llama-3.3-70b-instruct-pb6 8 NVFP4 No
llama-3.3-70b-instruct-pb6 1 NVFP4 Yes
llama-3.3-70b-instruct-pb6 2 NVFP4 Yes
llama-3.3-70b-instruct-pb6 4 NVFP4 Yes
llama-3.3-70b-instruct-pb6 8 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 1 BF16 No
llama-3.3-nemotron-super-49b-v1.5-pb6 2 BF16 No
llama-3.3-nemotron-super-49b-v1.5-pb6 4 BF16 No
llama-3.3-nemotron-super-49b-v1.5-pb6 8 BF16 No
llama-3.3-nemotron-super-49b-v1.5-pb6 1 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 2 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 4 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 8 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 1 FP8 No
llama-3.3-nemotron-super-49b-v1.5-pb6 2 FP8 No
llama-3.3-nemotron-super-49b-v1.5-pb6 4 FP8 No
llama-3.3-nemotron-super-49b-v1.5-pb6 8 FP8 No
llama-3.3-nemotron-super-49b-v1.5-pb6 1 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 2 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 4 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 8 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 1 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5-pb6 2 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5-pb6 4 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5-pb6 8 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5-pb6 1 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 2 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 4 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5-pb6 8 NVFP4 Yes
nemotron-3-nano-pb6 1 BF16 No
nemotron-3-nano-pb6 2 BF16 No
nemotron-3-nano-pb6 4 BF16 No
nemotron-3-nano-pb6 8 BF16 No
nemotron-3-nano-pb6 1 BF16 Yes
nemotron-3-nano-pb6 2 BF16 Yes
nemotron-3-nano-pb6 4 BF16 Yes
nemotron-3-nano-pb6 8 BF16 Yes
nemotron-3-nano-pb6 1 FP8 No
nemotron-3-nano-pb6 2 FP8 No
nemotron-3-nano-pb6 4 FP8 No
nemotron-3-nano-pb6 8 FP8 No
nemotron-3-nano-pb6 1 FP8 Yes
nemotron-3-nano-pb6 2 FP8 Yes
nemotron-3-nano-pb6 4 FP8 Yes
nemotron-3-nano-pb6 8 FP8 Yes
nemotron-3-nano-pb6 1 NVFP4 No
nemotron-3-nano-pb6 2 NVFP4 No
nemotron-3-nano-pb6 4 NVFP4 No
nemotron-3-nano-pb6 8 NVFP4 No
gpt-oss-120b-pb6 1 MXFP4 No
gpt-oss-120b-pb6 2 MXFP4 No
gpt-oss-120b-pb6 4 MXFP4 No
gpt-oss-120b-pb6 8 MXFP4 No
gpt-oss-120b-pb6 1 MXFP4 Yes
gpt-oss-120b-pb6 2 MXFP4 Yes
gpt-oss-120b-pb6 4 MXFP4 Yes
gpt-oss-120b-pb6 8 MXFP4 Yes
No matching profiles
Refer to Model-Free NIM

llama-3.1-8b-instruct-pb6#

Latest supported NIM LLM PB version: 2.0.4-pb6.3

The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct-pb6:

Precision

TP1

BF16

vllm-bf16-tp1-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

FP8

vllm-fp8-tp1-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-70b-instruct-pb6#

Latest supported NIM LLM PB version: 2.0.4-pb6.3

The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct-pb6:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A10G

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-nemotron-super-49b-v1.5-pb6#

Latest supported NIM LLM PB version: 2.0.4-pb6.3

The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5-pb6:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A10G

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-3-nano-pb6#

Latest supported NIM LLM PB version: 2.0.4-pb6.3

The following table lists the supported profile configurations for nvidia/nemotron-3-nano-pb6:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

gpt-oss-120b-pb6#

Latest supported NIM LLM PB version: 2.0.4-pb6.3

The following table lists the supported profile configurations for openai/gpt-oss-120b-pb6:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

model-free-nim-pb6#

Latest supported NIM LLM PB version: 2.0.4-pb6.3

The following models are tested and validated for nvidia/model-free-nim-pb6:

  • gpt-oss-20b

  • apriel-nemotron

  • codestral

While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.

Verified GPUs

The model-free NIM has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition