Support Matrix for NIMs#

This page lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM.

Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.

Use the table below to filter NIM profiles by GPU, tensor parallelism (TP), precision, LoRA support, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.

Note

The GB300-WS GPU is more commonly known as DGX Station.

ModelTPPrecisionLoRA
gpt-oss-120b 1 MXFP4 No
gpt-oss-120b 2 MXFP4 No
gpt-oss-120b 4 MXFP4 No
gpt-oss-120b 8 MXFP4 No
gpt-oss-120b 1 MXFP4 Yes
gpt-oss-120b 2 MXFP4 Yes
gpt-oss-120b 4 MXFP4 Yes
gpt-oss-120b 8 MXFP4 Yes
gpt-oss-20b 1 MXFP4 No
gpt-oss-20b 2 MXFP4 No
gpt-oss-20b 4 MXFP4 No
gpt-oss-20b 8 MXFP4 No
gpt-oss-20b 1 MXFP4 Yes
gpt-oss-20b 2 MXFP4 Yes
gpt-oss-20b 4 MXFP4 Yes
gpt-oss-20b 8 MXFP4 Yes
llama-3.1-70b-instruct 1 BF16 No
llama-3.1-70b-instruct 2 BF16 No
llama-3.1-70b-instruct 4 BF16 No
llama-3.1-70b-instruct 8 BF16 No
llama-3.1-70b-instruct 1 BF16 Yes
llama-3.1-70b-instruct 2 BF16 Yes
llama-3.1-70b-instruct 4 BF16 Yes
llama-3.1-70b-instruct 8 BF16 Yes
llama-3.1-70b-instruct 1 FP8 No
llama-3.1-70b-instruct 2 FP8 No
llama-3.1-70b-instruct 4 FP8 No
llama-3.1-70b-instruct 8 FP8 No
llama-3.1-70b-instruct 1 FP8 Yes
llama-3.1-70b-instruct 2 FP8 Yes
llama-3.1-70b-instruct 4 FP8 Yes
llama-3.1-70b-instruct 8 FP8 Yes
llama-3.1-70b-instruct 1 NVFP4 No
llama-3.1-70b-instruct 2 NVFP4 No
llama-3.1-70b-instruct 4 NVFP4 No
llama-3.1-70b-instruct 8 NVFP4 No
llama-3.1-70b-instruct 1 NVFP4 Yes
llama-3.1-70b-instruct 2 NVFP4 Yes
llama-3.1-70b-instruct 4 NVFP4 Yes
llama-3.1-70b-instruct 8 NVFP4 Yes
llama-3.1-8b-instruct 1 BF16 No
llama-3.1-8b-instruct 1 BF16 Yes
llama-3.1-8b-instruct 1 FP8 No
llama-3.1-8b-instruct 1 FP8 Yes
llama-3.1-8b-instruct 1 NVFP4 No
llama-3.1-8b-instruct 1 NVFP4 Yes
llama-3.3-70b-instruct 1 BF16 No
llama-3.3-70b-instruct 2 BF16 No
llama-3.3-70b-instruct 4 BF16 No
llama-3.3-70b-instruct 8 BF16 No
llama-3.3-70b-instruct 1 BF16 Yes
llama-3.3-70b-instruct 2 BF16 Yes
llama-3.3-70b-instruct 4 BF16 Yes
llama-3.3-70b-instruct 8 BF16 Yes
llama-3.3-70b-instruct 1 FP8 No
llama-3.3-70b-instruct 2 FP8 No
llama-3.3-70b-instruct 4 FP8 No
llama-3.3-70b-instruct 8 FP8 No
llama-3.3-70b-instruct 1 FP8 Yes
llama-3.3-70b-instruct 2 FP8 Yes
llama-3.3-70b-instruct 4 FP8 Yes
llama-3.3-70b-instruct 8 FP8 Yes
llama-3.3-70b-instruct 1 NVFP4 No
llama-3.3-70b-instruct 2 NVFP4 No
llama-3.3-70b-instruct 4 NVFP4 No
llama-3.3-70b-instruct 8 NVFP4 No
llama-3.3-70b-instruct 1 NVFP4 Yes
llama-3.3-70b-instruct 2 NVFP4 Yes
llama-3.3-70b-instruct 4 NVFP4 Yes
llama-3.3-70b-instruct 8 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 1 BF16 No
llama-3.3-nemotron-super-49b-v1.5 2 BF16 No
llama-3.3-nemotron-super-49b-v1.5 4 BF16 No
llama-3.3-nemotron-super-49b-v1.5 8 BF16 No
llama-3.3-nemotron-super-49b-v1.5 1 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 2 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 4 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 8 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 1 FP8 No
llama-3.3-nemotron-super-49b-v1.5 2 FP8 No
llama-3.3-nemotron-super-49b-v1.5 4 FP8 No
llama-3.3-nemotron-super-49b-v1.5 8 FP8 No
llama-3.3-nemotron-super-49b-v1.5 1 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 2 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 4 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 8 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 1 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 2 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 4 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 8 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 1 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 2 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 4 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 8 NVFP4 Yes
nemotron-content-safety-reasoning-4b-experimental 1 BF16 No
nemotron-content-safety-reasoning-4b-experimental 2 BF16 No
nemotron-content-safety-reasoning-4b-experimental 4 BF16 No
nemotron-content-safety-reasoning-4b-experimental 8 BF16 No
nemotron-3-nano 1 BF16 No
nemotron-3-nano 2 BF16 No
nemotron-3-nano 4 BF16 No
nemotron-3-nano 8 BF16 No
nemotron-3-nano 1 BF16 Yes
nemotron-3-nano 2 BF16 Yes
nemotron-3-nano 4 BF16 Yes
nemotron-3-nano 8 BF16 Yes
nemotron-3-nano 1 FP8 No
nemotron-3-nano 2 FP8 No
nemotron-3-nano 4 FP8 No
nemotron-3-nano 8 FP8 No
nemotron-3-nano 1 NVFP4 No
nemotron-3-nano 2 NVFP4 No
nemotron-3-nano 4 NVFP4 No
nemotron-3-nano 8 NVFP4 No
nemotron-3-nano 1 NVFP4 Yes
nemotron-3-super-120b-a12b 1 BF16 No
nemotron-3-super-120b-a12b 2 BF16 No
nemotron-3-super-120b-a12b 4 BF16 No
nemotron-3-super-120b-a12b 8 BF16 No
nemotron-3-super-120b-a12b 2 BF16 Yes
nemotron-3-super-120b-a12b 4 BF16 Yes
nemotron-3-super-120b-a12b 8 BF16 Yes
nemotron-3-super-120b-a12b 1 FP8 No
nemotron-3-super-120b-a12b 2 FP8 No
nemotron-3-super-120b-a12b 4 FP8 No
nemotron-3-super-120b-a12b 8 FP8 No
nemotron-3-super-120b-a12b 1 NVFP4 No
nemotron-3-super-120b-a12b 2 NVFP4 No
nemotron-3-super-120b-a12b 4 NVFP4 No
nemotron-3-super-120b-a12b 8 NVFP4 No
nemotron-3-super-120b-a12b 1 NVFP4 Yes
nemotron-3-super-120b-a12b 2 NVFP4 Yes
nemotron-3-ultra-550b-a55b 8 BF16 No
nemotron-3-ultra-550b-a55b 8 BF16 Yes
nemotron-3-ultra-550b-a55b 2 NVFP4 No
nemotron-3-ultra-550b-a55b 4 NVFP4 No
nemotron-3-ultra-550b-a55b 8 NVFP4 No
nemotron-3.5-lightning-30b-a3b 1 BF16 No
nemotron-3.5-lightning-30b-a3b 1 BF16 Yes
nemotron-3.5-lightning-30b-a3b 2 BF16 No
nemotron-3.5-lightning-30b-a3b 2 BF16 Yes
nemotron-3.5-lightning-30b-a3b 4 BF16 No
nemotron-3.5-lightning-30b-a3b 4 BF16 Yes
nemotron-3.5-lightning-30b-a3b 8 BF16 No
nemotron-3.5-lightning-30b-a3b 8 BF16 Yes
nemotron-3.5-lightning-30b-a3b 1 W4A16 No
nemotron-3.5-lightning-30b-a3b 1 W4A16 Yes
nemotron-3.5-lightning-30b-a3b 2 W4A16 No
nemotron-3.5-lightning-30b-a3b 2 W4A16 Yes
nemotron-3.5-lightning-30b-a3b 4 W4A16 No
nemotron-3.5-lightning-30b-a3b 4 W4A16 Yes
nemotron-3.5-lightning-30b-a3b 1 NVFP4 No
nemotron-3.5-lightning-30b-a3b 1 NVFP4 Yes
nemotron-3.5-lightning-30b-a3b 2 NVFP4 No
nemotron-3.5-lightning-30b-a3b 2 NVFP4 Yes
nemotron-3.5-lightning-30b-a3b 4 NVFP4 No
nemotron-3.5-lightning-30b-a3b 4 NVFP4 Yes
riva-translate-4b-instruct-v2 1 BF16 No
starcoder2-7b 1 BF16 No
starcoder2-7b 2 BF16 No

gpt-oss-120b#

The following table lists the supported profile configurations for openai/gpt-oss-120b:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

gpt-oss-20b#

The following table lists the supported profile configurations for openai/gpt-oss-20b:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.1-70b-instruct#

The following table lists the supported profile configurations for meta/llama-3.1-70b-instruct:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.1-8b-instruct#

The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct:

Precision

TP1

BF16

vllm-bf16-tp1-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

FP8

vllm-fp8-tp1-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-70b-instruct#

The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-nemotron-super-49b-v1.5#

The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-content-safety-reasoning-4b-experimental#

The following table lists the supported profile configurations for nvidia/nemotron-content-safety-reasoning-4b-experimental:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

For information on getting started with nemotron-content-safety-reasoning-4b-experimental, refer to version 2.0.3 of the documentation.

nemotron-3-nano#

The following table lists the supported profile configurations for nvidia/nemotron-3-nano:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-3-super-120b-a12b#

The following table lists the supported profile configurations for nvidia/nemotron-3-super-120b-a12b:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

Note

This is a large model. Lower-TP profiles require substantially more GPU memory per device, so some verified GPUs support only TP4 or TP8 profiles.

Verified GPUs

The following GPUs have been verified with one or more supported profiles for this model:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-3-ultra-550b-a55b#

The following table lists the supported profile configurations for nvidia/nemotron-3-ultra-550b-a55b:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

Note

The minimum deployment is two GPUs, using vllm-nvfp4-tp2-pp1 on high-memory Blackwell SKUs (NVIDIA-B300-SXM6-AC and NVIDIA-GB300); the other profiles use four or eight GPUs. H200 validation covers one NVFP4 TP4 profile and is less extensive than the Blackwell validation.

FP8 profiles were not released for this NIM because NVFP4 profiles work on all tested Hopper-architecture GPUs (for example, H100 and H200) and deliver better performance than FP8 profiles.

Verified GPUs

The following GPUs have been verified with one or more supported profiles for this model:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

nemotron-3.5-lightning-30b-a3b#

The following table lists the supported profile configurations for nvidia/nemotron-3.5-lightning-30b-a3b, along with the minimum GPU VRAM required per device and the GPU architecture that each precision requires. The VRAM values come from the profile disk-sizing metadata reported by list-model-profiles and represent the minimum per-GPU memory needed to load the model weights and the runtime allocations; they do not include additional headroom for large KV caches at high context length or high concurrency. Refer to Resolve Out-of-Memory Errors for headroom guidance.

Nemotron 3.5 Lightning Profile Information#

Precision

TP

PP

Min VRAM per GPU

Min GPU count

Architecture requirement

LoRA Profile

BF16

1

1

66 GB

1

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

2

1

35 GB

2

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

4

1

20 GB

4

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

8

1

12 GB

8

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

1

1

32 GB

1

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

2

1

20 GB

2

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

4

1

14 GB

4

Ampere or newer (SM 8.0+)

Base and LoRA

NVFP4

1

1

30 GB

1

Blackwell or newer (SM 10.0+)

Base and LoRA

NVFP4

2

1

18 GB

2

Blackwell or newer (SM 10.0+)

Base and LoRA

NVFP4

4

1

12 GB

4

Blackwell or newer (SM 10.0+)

Base and LoRA

Notes:

  • Min VRAM per GPU is the floor that list-model-profiles reports for that precision and tensor-parallel size. A GPU below the floor is filtered out during profile selection.

  • Architecture requirement reflects the kernels each precision needs. NVFP4 requires Streaming Multiprocessor (SM) 10.0+ (Blackwell) because it uses NVFP4-native tensor cores. BF16 and W4A16 run on any Ampere-class or newer GPU (SM 8.0+).

  • Sharing weights across more GPUs (higher TP) lowers the per-GPU VRAM floor and enables lower-memory GPUs.

  • Model-cache disk sizes range from approximately 19 GB (NVFP4) to approximately 63 GB (BF16). Cache size is separate from GPU VRAM and does not indicate that the model will fit in a given GPU.

Verified GPUs

This model has been verified on the following GPUs (grouped by architecture):

Blackwell (SM 10.0) — supports NVFP4, W4A16, BF16:

GPU

VRAM per GPU

Profiles that fit

NVIDIA-B200

180 GB

All profiles

NVIDIA-B300-SXM6-AC

288 GB

All profiles

NVIDIA-GB200

186 GB

All profiles

NVIDIA-GB300

288 GB

All profiles

NVIDIA-GB10

128 GB (unified memory)

All profiles; use --gpu-memory-utilization 0.75 on GB10 (see the getting started guide)

NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

96 GB

All profiles

Hopper (SM 9.0) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-H100-80GB-HBM3

80 GB

BF16 all sizes (needs --max-num-seqs 512 on TP=1), W4A16 all sizes

NVIDIA-H200

141 GB

BF16 all sizes, W4A16 all sizes

Ampere (SM 8.0) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-A100-SXM4-80GB

80 GB

BF16 TP≥2, W4A16 all sizes

Ada (SM 8.9) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-L40S

48 GB

BF16 TP≥2, W4A16 all sizes

Unverified but potentially compatible. Any GPU that meets the per-profile minimum VRAM floor and the architecture requirement in the profile table can run that profile technically, even if it is not in the preceding Verified GPUs tables. Only the GPUs in those tables have been tested end-to-end for this model at this NIM version; untested systems may exhibit different startup, throughput, or accuracy behavior.

riva-translate-4b-instruct-v2#

The following table lists the supported profile configuration for nvidia/riva-translate-4b-instruct-v2:

Precision

TP

PP

LoRA

Profile

BF16

1

1

No

vllm-bf16-tp1-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

starcoder2-7b#

The following table lists the supported profile configurations for bigcode/starcoder2-7b:

Precision

TP1

TP2

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-80GB-PCIe

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-GB300-WS

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

Model-Free NIM#

The following models are tested and validated for nvidia/model-free-nim:

  • gpt-oss-20b

  • apriel-nemotron

  • codestral

While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.

Verified GPUs

The model-free NIM has been verified on the following GPUs:

  • NVIDIA-A100-80GB-PCIe

  • NVIDIA-A100-PCIE-40GB

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H100-PCIe

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

1.x NIM LLM Models#

For more information on version 1.x NIMs, refer to the 1.15 version of the NIM LLM Supported Models page.

Show 1.x models