Support Matrix for NIMs#

This page lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM and VLM. This page is organized according to model type:

Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.

Optimized profiles are listed separately in the model sections because multiple configurations can share the same precision and tensor-parallel size. For selection instructions, refer to Model Profiles and Selection.

Important

The list of NIM Certified models is available on the NGC Catalog site.

Use the table below to filter NIM profiles by GPU, tensor parallelism (TP), precision, model type, LoRA support, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.

Optimized profile entries identify their tuning objective and link to the full profile ID in the model section. The filters also apply to rows in the optimized profile tables below.

Note

The GB300-WS GPU is more commonly known as DGX Station.

ModelTPPrecisionLoRA
deepseek-v4-pro-0813 8 FP8 No
gemma-4-26b-a4b-it 1 BF16 No
gemma-4-26b-a4b-it 2 BF16 No
gemma-4-26b-a4b-it 4 BF16 No
gemma-4-26b-a4b-it 1 NVFP4 No
gemma-4-31b-it 1 BF16 No
gemma-4-31b-it 2 BF16 No
gemma-4-31b-it 4 BF16 No
gemma-4-31b-it 2 INT4 No
gemma-4-31b-it 1 NVFP4 No
gemma-4-31b-it 2 NVFP4 No
gemma-4-31b-it 4 NVFP4 No
gpt-oss-120b 1 MXFP4 No
gpt-oss-120b 2 MXFP4 No
gpt-oss-120b 4 MXFP4 No
gpt-oss-120b 8 MXFP4 No
gpt-oss-120b 1 MXFP4 Yes
gpt-oss-120b 2 MXFP4 Yes
gpt-oss-120b 4 MXFP4 Yes
gpt-oss-120b 8 MXFP4 Yes
gpt-oss-20b 1 MXFP4 No
gpt-oss-20b 2 MXFP4 No
gpt-oss-20b 4 MXFP4 No
gpt-oss-20b 8 MXFP4 No
gpt-oss-20b 1 MXFP4 Yes
gpt-oss-20b 2 MXFP4 Yes
gpt-oss-20b 4 MXFP4 Yes
gpt-oss-20b 8 MXFP4 Yes
kimi-k2.6 4 NVFP4 No
kimi-k2.6 8 INT4 No
llama-3.1-70b-instruct 1 BF16 No
llama-3.1-70b-instruct 2 BF16 No
llama-3.1-70b-instruct 4 BF16 No
llama-3.1-70b-instruct 8 BF16 No
llama-3.1-70b-instruct 1 BF16 Yes
llama-3.1-70b-instruct 2 BF16 Yes
llama-3.1-70b-instruct 4 BF16 Yes
llama-3.1-70b-instruct 8 BF16 Yes
llama-3.1-70b-instruct 1 FP8 No
llama-3.1-70b-instruct 2 FP8 No
llama-3.1-70b-instruct 4 FP8 No
llama-3.1-70b-instruct 8 FP8 No
llama-3.1-70b-instruct 1 FP8 Yes
llama-3.1-70b-instruct 2 FP8 Yes
llama-3.1-70b-instruct 4 FP8 Yes
llama-3.1-70b-instruct 8 FP8 Yes
llama-3.1-70b-instruct 1 NVFP4 No
llama-3.1-70b-instruct 2 NVFP4 No
llama-3.1-70b-instruct 4 NVFP4 No
llama-3.1-70b-instruct 8 NVFP4 No
llama-3.1-70b-instruct 1 NVFP4 Yes
llama-3.1-70b-instruct 2 NVFP4 Yes
llama-3.1-70b-instruct 4 NVFP4 Yes
llama-3.1-70b-instruct 8 NVFP4 Yes
llama-3.1-8b-instruct 1 BF16 No
llama-3.1-8b-instruct 1 BF16 Yes
llama-3.1-8b-instruct 1 FP8 No
llama-3.1-8b-instruct 1 FP8 Yes
llama-3.1-8b-instruct 1 NVFP4 No
llama-3.1-8b-instruct 1 NVFP4 Yes
llama-3.3-70b-instruct 1 BF16 No
llama-3.3-70b-instruct 2 BF16 No
llama-3.3-70b-instruct 4 BF16 No
llama-3.3-70b-instruct 8 BF16 No
llama-3.3-70b-instruct 1 BF16 Yes
llama-3.3-70b-instruct 2 BF16 Yes
llama-3.3-70b-instruct 4 BF16 Yes
llama-3.3-70b-instruct 8 BF16 Yes
llama-3.3-70b-instruct 1 FP8 No
llama-3.3-70b-instruct 2 FP8 No
llama-3.3-70b-instruct 4 FP8 No
llama-3.3-70b-instruct 8 FP8 No
llama-3.3-70b-instruct 1 FP8 Yes
llama-3.3-70b-instruct 2 FP8 Yes
llama-3.3-70b-instruct 4 FP8 Yes
llama-3.3-70b-instruct 8 FP8 Yes
llama-3.3-70b-instruct 1 NVFP4 No
llama-3.3-70b-instruct 2 NVFP4 No
llama-3.3-70b-instruct 4 NVFP4 No
llama-3.3-70b-instruct 8 NVFP4 No
llama-3.3-70b-instruct 1 NVFP4 Yes
llama-3.3-70b-instruct 2 NVFP4 Yes
llama-3.3-70b-instruct 4 NVFP4 Yes
llama-3.3-70b-instruct 8 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 1 BF16 No
llama-3.3-nemotron-super-49b-v1.5 2 BF16 No
llama-3.3-nemotron-super-49b-v1.5 4 BF16 No
llama-3.3-nemotron-super-49b-v1.5 8 BF16 No
llama-3.3-nemotron-super-49b-v1.5 1 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 2 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 4 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 8 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 1 FP8 No
llama-3.3-nemotron-super-49b-v1.5 2 FP8 No
llama-3.3-nemotron-super-49b-v1.5 4 FP8 No
llama-3.3-nemotron-super-49b-v1.5 8 FP8 No
llama-3.3-nemotron-super-49b-v1.5 1 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 2 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 4 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 8 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 1 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 2 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 4 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 8 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 1 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 2 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 4 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 8 NVFP4 Yes
mistral-small-4-119b-2603 1 FP8 No
mistral-small-4-119b-2603 2 FP8 No
mistral-small-4-119b-2603 4 FP8 No
mistral-small-4-119b-2603 8 FP8 No
mistral-small-4-119b-2603 1 NVFP4 No
mistral-small-4-119b-2603 2 NVFP4 No
mistral-small-4-119b-2603 4 NVFP4 No
mistral-small-4-119b-2603 8 NVFP4 No
nemotron-content-safety-reasoning-4b-experimental 1 BF16 No
nemotron-content-safety-reasoning-4b-experimental 2 BF16 No
nemotron-content-safety-reasoning-4b-experimental 4 BF16 No
nemotron-content-safety-reasoning-4b-experimental 8 BF16 No
nemotron-3-nano 1 BF16 No
nemotron-3-nano 2 BF16 No
nemotron-3-nano 4 BF16 No
nemotron-3-nano 8 BF16 No
nemotron-3-nano 1 BF16 Yes
nemotron-3-nano 2 BF16 Yes
nemotron-3-nano 4 BF16 Yes
nemotron-3-nano 8 BF16 Yes
nemotron-3-nano 1 FP8 No
nemotron-3-nano 2 FP8 No
nemotron-3-nano 4 FP8 No
nemotron-3-nano 8 FP8 No
nemotron-3-nano 1 NVFP4 No
nemotron-3-nano 2 NVFP4 No
nemotron-3-nano 4 NVFP4 No
nemotron-3-nano 8 NVFP4 No
nemotron-3-nano 1 NVFP4 Yes
nemotron-3-nano-omni-30b-a3b-reasoning 2 BF16 No
nemotron-3-nano-omni-30b-a3b-reasoning 1 FP8 No
nemotron-3-nano-omni-30b-a3b-reasoning 1 NVFP4 No
nemotron-3-super-120b-a12b 1 BF16 No
nemotron-3-super-120b-a12b 2 BF16 No
nemotron-3-super-120b-a12b 4 BF16 No
nemotron-3-super-120b-a12b 8 BF16 No
nemotron-3-super-120b-a12b 2 BF16 Yes
nemotron-3-super-120b-a12b 4 BF16 Yes
nemotron-3-super-120b-a12b 8 BF16 Yes
nemotron-3-super-120b-a12b 1 FP8 No
nemotron-3-super-120b-a12b 2 FP8 No
nemotron-3-super-120b-a12b 4 FP8 No
nemotron-3-super-120b-a12b 8 FP8 No
nemotron-3-super-120b-a12b 1 NVFP4 No
nemotron-3-super-120b-a12b 2 NVFP4 No
nemotron-3-super-120b-a12b 4 NVFP4 No
nemotron-3-super-120b-a12b 8 NVFP4 No
nemotron-3-super-120b-a12b 1 NVFP4 Yes
nemotron-3-super-120b-a12b 2 NVFP4 Yes
nemotron-3-ultra-550b-a55b 8 BF16 No
nemotron-3-ultra-550b-a55b 2 NVFP4 No
nemotron-3-ultra-550b-a55b 4 NVFP4 No
nemotron-3-ultra-550b-a55b 8 NVFP4 No
nemotron-3.5-lightning 1 BF16 No
nemotron-3.5-lightning 1 BF16 Yes
nemotron-3.5-lightning 2 BF16 No
nemotron-3.5-lightning 2 BF16 Yes
nemotron-3.5-lightning 4 BF16 No
nemotron-3.5-lightning 4 BF16 Yes
nemotron-3.5-lightning 8 BF16 No
nemotron-3.5-lightning 8 BF16 Yes
nemotron-3.5-lightning 1 W4A16 No
nemotron-3.5-lightning 1 W4A16 Yes
nemotron-3.5-lightning 2 W4A16 No
nemotron-3.5-lightning 2 W4A16 Yes
nemotron-3.5-lightning 4 W4A16 No
nemotron-3.5-lightning 4 W4A16 Yes
nemotron-3.5-lightning 1 NVFP4 No
nemotron-3.5-lightning 1 NVFP4 Yes
nemotron-3.5-lightning 2 NVFP4 No
nemotron-3.5-lightning 2 NVFP4 Yes
nemotron-3.5-lightning 4 NVFP4 No
nemotron-3.5-lightning 4 NVFP4 Yes
qwen3.5-122b-a10b 1 BF16 No
qwen3.5-122b-a10b 1 BF16 Yes
qwen3.5-122b-a10b 2 BF16 No
qwen3.5-122b-a10b 2 BF16 Yes
qwen3.5-122b-a10b 4 BF16 No
qwen3.5-122b-a10b 4 BF16 Yes
qwen3.5-122b-a10b 8 BF16 No
qwen3.5-122b-a10b 8 BF16 Yes
qwen3.5-122b-a10b 1 FP8 No
qwen3.5-122b-a10b 1 FP8 Yes
qwen3.5-122b-a10b 2 FP8 No
qwen3.5-122b-a10b 2 FP8 Yes
qwen3.5-122b-a10b 4 FP8 No
qwen3.5-122b-a10b 4 FP8 Yes
qwen3.5-122b-a10b 8 FP8 No
qwen3.5-122b-a10b 8 FP8 Yes
qwen3.5-122b-a10b 1 NVFP4 No
qwen3.5-122b-a10b 1 NVFP4 Yes
qwen3.5-122b-a10b 2 NVFP4 No
qwen3.5-122b-a10b 2 NVFP4 Yes
qwen3.5-122b-a10b 4 NVFP4 No
qwen3.5-122b-a10b 4 NVFP4 Yes
qwen3.5-122b-a10b 8 NVFP4 No
qwen3.5-122b-a10b 8 NVFP4 Yes
qwen3.5-397b-a17b 4 BF16 No
qwen3.5-397b-a17b 4 BF16 Yes
qwen3.5-397b-a17b 8 BF16 No
qwen3.5-397b-a17b 8 BF16 Yes
qwen3.5-397b-a17b 2 FP8 No
qwen3.5-397b-a17b 2 FP8 Yes
qwen3.5-397b-a17b 4 FP8 No
qwen3.5-397b-a17b 4 FP8 Yes
qwen3.5-397b-a17b 8 FP8 No
qwen3.5-397b-a17b 8 FP8 Yes
qwen3.5-397b-a17b 1 NVFP4 No
qwen3.5-397b-a17b 1 NVFP4 Yes
qwen3.5-397b-a17b 2 NVFP4 No
qwen3.5-397b-a17b 2 NVFP4 Yes
qwen3.5-397b-a17b 4 NVFP4 No
qwen3.5-397b-a17b 4 NVFP4 Yes
qwen3.5-397b-a17b 8 NVFP4 No
qwen3.5-397b-a17b 8 NVFP4 Yes
riva-translate-4b-instruct-v2 1 BF16 No
starcoder2-7b 1 BF16 No
starcoder2-7b 2 BF16 No

LLMs#

The following section lists all of the supported LLMs.

deepseek-v4-pro-0813#

The following table lists the supported profile configurations for deepseek-ai/DeepSeek-V4-Pro-0813.

This NIM is an SGLang 2.1.2-variant container. Combined container and model disk space is approximately 831 GB. The fallback profile requires more than 1,024 GB of aggregate GPU memory across eight GPUs. For more information, refer to the dedicated Get Started guide.

GPU

GPU Memory

Precision

TP

PP

LoRA

Profile

Any (fallback)

> 1,024 GB aggregate

FP8

8

1

No

sglang-fp8-tp8-fallback-1024g

H20-3e

141 GB

FP8

8

1

No

sglang-h20-3e-fp8-tp8-throughput-232c:10de-1024g

H200

141 GB

FP8

8

1

No

sglang-h200-fp8-tp8-throughput-2335:10de-1024g

B200

196 GB

FP8

8

1

No

sglang-b200-fp8-tp8-throughput-2901:10de-1024g

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-H20-3e

  • NVIDIA-H200

gpt-oss-120b#

The following table lists the supported profile configurations for openai/gpt-oss-120b:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

MXFP4

1

NVIDIA-B200, NVIDIA-B300-SXM6-AC, NVIDIA-GB200, NVIDIA-GB300, NVIDIA-GB300-WS, NVIDIA-GH200-144G-HBM3e, NVIDIA-GH200-480GB, NVIDIA-H100-80GB-HBM3, NVIDIA-H100-NVL, NVIDIA-H200, NVIDIA-H200-NVL, NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

No

Throughput

7de9349cde0a225c2a05f727fbabb8192c067a6d3b1440051730dd9f5c078388

gpt-oss-20b#

The following table lists the supported profile configurations for openai/gpt-oss-20b:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.1-70b-instruct#

The following table lists the supported profile configurations for meta/llama-3.1-70b-instruct:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

2

NVIDIA-B200

No

Throughput

e1d3ae3b1367b98b9a7a729f3087d8e411e709f436c624129b939f96dcfd02b7

FP8

2

NVIDIA-H200

No

Throughput

421adb20111d74d64424eab4d8e9ed1a4cbd756338f03c52a32c4ee955301044

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.1-8b-instruct#

The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct:

Precision

TP1

BF16

vllm-bf16-tp1-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

FP8

vllm-fp8-tp1-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

381790f97c25437119051f987f5b172b00e6e5281545c76e5318caa9880ecd15

FP8

1

NVIDIA-H200

No

Throughput

86e27d488d1a6a6b7fa5fb2a6894bafc7059376021286fab43fcd9c612abd7f4

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-70b-instruct#

The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

16d3616335859278752b8ec6bdd3d88657880120895118fcd584c75f6f3249e0

FP8

2

NVIDIA-H200

No

Throughput

f8197e13fcb703fc7df8a9060d3d2be9745b4ff98b0eac3e466779afc127f2a6

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-nemotron-super-49b-v1.5#

The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

d6eb0398f20918bf0f1ee7703db82d00c5af0580f21facc8ad2a40d1c82c0ba7

FP8

2

NVIDIA-H200

No

Throughput

9923047c5f78eca58a4a1d586e85cfaaa4c724d5679bb615a9e954b4530033c3

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-content-safety-reasoning-4b-experimental#

The following table lists the supported profile configurations for nvidia/nemotron-content-safety-reasoning-4b-experimental. The latest release version of this NIM is 2.0.3. For more information, refer to the dedicated Get Started guide.

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

nemotron-3-nano#

The following table lists the supported profile configurations for nvidia/nemotron-3-nano:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

–

–

–

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

669d73c7ab89d5b53b392716cbcd8a0e06cf40d9d5051aae0b9f9a52891933a9

FP8

1

NVIDIA-H200

No

Throughput

8ca069b289f1838b0c63aeb51317f623f390975c099ba57f805c6c8940a81dad

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-3-super-120b-a12b#

The following table lists the supported profile configurations for nvidia/nemotron-3-super-120b-a12b:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

–

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

–

–

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

4

NVIDIA-B200

No

Throughput

3c7c3d74b044f7cf4bba1a80f3f596405a17efa35c663ca52963d71aaf381ea3

FP8

4

NVIDIA-H200

No

Throughput

ca31277ebd3b387e1a597a550541ab9b8dde3b31f6b5b5695ec8799c0162573c

Note

This is a large model. Lower-TP profiles require substantially more GPU memory per device, so some verified GPUs support only TP4 or TP8 profiles.

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-3-ultra-550b-a55b#

The following table lists the supported profile configurations for nvidia/nemotron-3-ultra-550b-a55b:

Precision

TP1

TP2

TP4

TP8

BF16

–

–

–

vllm-bf16-tp8-pp1

BF16 + LoRA

–

–

–

–

NVFP4

–

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

4

NVIDIA-B200

No

Throughput

7cdfe4d76f5fd32ad9ae240eef8b5d4af81dcc0fa23c6f20c4bc7da1a389a92c

NVFP4

4

NVIDIA-B300-SXM6-AC

No

Throughput

5b3441c9d0f55e8b4442537a4294304d5401d3a5542870bda16f3f8e18971878

NVFP4

4

NVIDIA-GB200

No

Throughput

a953e969777d28fc50e322d31a1bdfaf6aeb1a3cdff63ca3760cf0d424b9a90a

NVFP4

4

NVIDIA-GB300

No

Throughput

4466f9f08faedae5362260c2552af473240428fd9bf17d7bbaaf5a535abdde1e

Note

The minimum deployment is two GPUs, using vllm-nvfp4-tp2-pp1 on high-memory Blackwell SKUs (NVIDIA-B300-SXM6-AC and NVIDIA-GB300); the other profiles use four or eight GPUs. H200 validation covers one NVFP4 TP4 profile and is less extensive than the Blackwell validation.

FP8 profiles were not released for this NIM because NVFP4 profiles work on all tested Hopper-architecture GPUs (for example, H100 and H200) and deliver better performance than FP8 profiles.

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

nemotron-3.5-lightning#

The following table lists the supported profile configurations for nvidia/nemotron-3.5-lightning, along with the minimum GPU VRAM required per device and the GPU architecture that each precision requires. The VRAM values come from the profile disk-sizing metadata reported by list-model-profiles and represent the minimum per-GPU memory needed to load the model weights and the runtime allocations; they do not include additional headroom for large KV caches at high context length or high concurrency.

Nemotron 3.5 Lightning Profile Information#

Precision

TP

PP

Min VRAM per GPU

Min GPU count

Architecture requirement

LoRA Profile

BF16

1

1

66 GB

1

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

2

1

35 GB

2

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

4

1

20 GB

4

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

8

1

12 GB

8

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

1

1

32 GB

1

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

2

1

20 GB

2

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

4

1

14 GB

4

Ampere or newer (SM 8.0+)

Base and LoRA

NVFP4

1

1

30 GB

1

Blackwell or newer (SM 10.0+)

Base and LoRA

NVFP4

2

1

18 GB

2

Blackwell or newer (SM 10.0+)

Base and LoRA

NVFP4

4

1

12 GB

4

Blackwell or newer (SM 10.0+)

Base and LoRA

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

4816da181f28c435aaa2efa0e2302309c28f2c08c9534070e89f67854a521d72

NVFP4 (W4A16) + FP8

1

NVIDIA-H200

No

Throughput

1e34732516864e06c8dad271edd1d19b986c4c010b9d7e3adcb8b113df58819b

Notes:

  • Min VRAM per GPU is the floor that list-model-profiles reports for that precision and tensor-parallel size. A GPU below the floor is filtered out during profile selection.

  • Architecture requirement in the generic profile table reflects the kernels those profiles need. The generic NVFP4 profiles require Streaming Multiprocessor (SM) 10.0+ (Blackwell) because they use NVFP4-native tensor cores. BF16 and W4A16 run on any Ampere-class or newer GPU (SM 8.0+).

  • Sharing weights across more GPUs (higher TP) lowers the per-GPU VRAM floor and enables lower-memory GPUs.

  • Model-cache disk sizes range from approximately 19 GB (NVFP4) to approximately 63 GB (BF16). Cache size is separate from GPU VRAM and does not indicate that the model will fit in a given GPU.

Verified GPUs

The following tables list verified GPUs for the generic profiles, grouped by architecture:

Blackwell (SM 10.0) — supports NVFP4, W4A16, BF16:

GPU

VRAM per GPU

Profiles that fit

NVIDIA-B200

180 GB

All profiles

NVIDIA-B300-SXM6-AC

288 GB

All profiles

NVIDIA-GB200

186 GB

All profiles

NVIDIA-GB300

288 GB

All profiles

NVIDIA-GB10

128 GB (unified memory)

All profiles; use --gpu-memory-utilization 0.75 on GB10 (see the getting started guide)

NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

96 GB

All profiles

Hopper (SM 9.0) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-H100-80GB-HBM3

80 GB

BF16 all sizes (needs --max-num-seqs 512 on TP=1), W4A16 all sizes

NVIDIA-H200

141 GB

BF16 all sizes, W4A16 all sizes

Ampere (SM 8.0) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-A100-SXM4-80GB

80 GB

BF16 TP≥2, W4A16 all sizes

Ada (SM 8.9) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-L40S

48 GB

BF16 TP≥2, W4A16 all sizes

Unverified but potentially compatible. Any GPU that meets the per-profile minimum VRAM floor and the architecture requirement in the profile table can run that profile technically, even if it is not in the preceding Verified GPUs tables. Only the GPUs in those tables have been tested end-to-end for this model at this NIM version; untested systems may exhibit different startup, throughput, or accuracy behavior.

riva-translate-4b-instruct-v2#

The following table lists the supported profile configuration for nvidia/riva-translate-4b-instruct-v2.

The latest release version of this NIM is 2.0.8.

Precision

TP

PP

LoRA

Profile

BF16

1

1

No

vllm-bf16-tp1-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

starcoder2-7b#

The following table lists the supported profile configurations for bigcode/starcoder2-7b:

Precision

TP1

TP2

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

BF16

1

NVIDIA-A100-80GB-PCIe, NVIDIA-A100-SXM4-80GB, NVIDIA-B200, NVIDIA-GB300-WS, NVIDIA-H100-80GB-HBM3, NVIDIA-H200

No

Throughput

65053e7510a772fa5ebc8da386b9b593152c746173b4c2fb85a1739a276e0d2c

BF16

2

NVIDIA-A100-SXM4-80GB, NVIDIA-B200, NVIDIA-H100-80GB-HBM3, NVIDIA-H200

No

Throughput

39725ead141c0eb26d6cbdcce32660a6da469663d88bc7ea7441808ff0134ca5

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-80GB-PCIe

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-GB300-WS

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

VLMs#

The following section lists all of the supported VLMs.

gemma-4-26b-a4b-it#

The following table lists the supported profile configurations for google/gemma-4-26b-a4b-it:

Precision

TP1

TP2

TP4

TP8

NVFP4

vllm-nvfp4-tp1-pp1

–

–

–

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

–

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

234e46324452f6124bf432eaa99cc1221ed917b069ee63ebc0c4d2e636f22f3d

NVFP4

1

NVIDIA-GB200

No

Throughput

04094afda5cf6fc9df04560f322cc677678a0d4caf943d058ff1c9779aaf750c

BF16

4

NVIDIA-A10G

No

Context length

800f3c24e5bdaca88aebd04242be7765c65ac7abbfc9013c404d0ea0ddc9f6fc

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

gemma-4-31b-it#

The following table lists the supported profile configurations for google/gemma-4-31b-it:

Precision

TP1

TP2

TP4

TP8

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

–

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

–

INT4

–

vllm-int4-tp2-pp1

–

–

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200, NVIDIA-B300-SXM6-AC, NVIDIA-GB200, NVIDIA-GB300, NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

No

Throughput

5f2613453901ce34041f9390d1eb035bceabba0af65af73be562fa5cda704a48

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

kimi-k2.6#

The following table lists the supported profile configurations for moonshotai/kimi-k2.6:

Precision

TP1

TP2

TP4

TP8

NVFP4

–

–

vllm-nvfp4-tp4-pp1

–

INT4

–

–

–

vllm-int4-tp8-pp1

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

4

NVIDIA-B200

No

Throughput

918ca5e8acff6936981f6948e7b4d149bcd84eeb59a0eda9b70f0d18c55c8eab

NVFP4

4

NVIDIA-B300-SXM6-AC

No

Throughput

7e73d8d993b39c8bac790c0d2872658ac85d2b86227c1ae7c1393c5e63b3dab7

NVFP4

4

NVIDIA-GB200

No

Throughput

c1d37f672564cec2e2cd881a1da9e3a01ebc9134fdfb274fd09b00c93dfa770d

NVFP4

4

NVIDIA-GB300

No

Throughput

24bfc505637e8cad50ad862666b8e6769a08e9c029528fb1653711de05045987

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H200

mistral-small-4-119b-2603#

The following table lists the supported profile configurations for mistralai/mistral-small-4-119b-2603:

Precision

TP1

TP2

TP4

TP8

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H200-NVL

nemotron-3-nano-omni-30b-a3b-reasoning#

The following table lists the supported profile configurations for nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:

Precision

TP1

TP2

TP4

TP8

BF16

–

vllm-bf16-tp2-pp1

–

–

FP8

vllm-fp8-tp1-pp1

–

–

–

NVFP4

vllm-nvfp4-tp1-pp1

–

–

–

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

2883a35a92ca200be25ce99fa4f04c4807a37c3d7b01f2192e00ddd5f9d105b7

NVFP4

1

NVIDIA-B300-SXM6-AC

No

Throughput

7cb9d0358baa2792b09a9907acc032df060acd912ce9c5b99c99589e4f34182c

NVFP4

1

NVIDIA-GB300

No

Throughput

22da479a1490835d7980611878ca91da3dd8bfa8c63c72c814c8867955d9eff8

FP8

1

NVIDIA-H200

No

Throughput

babd305818fe475ed855900b0518a7ef2d6bdf3ccd651584094a706dd010822e

NVFP4

1

NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

No

Throughput

ffb4dc6902e6bb4c27948ea1700707c664b179ab214072f9f700b0fffdefa1de

NVFP4

1

NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

No

Throughput

d8e8c64a19e5150344259745878af4dc1eabc270326a8b7bb4516c90ec602e5c

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

qwen3.5-122b-a10b#

The following table lists the supported profile configurations for qwen/qwen3.5-122b-a10b:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

1

NVIDIA-B200

No

Throughput

5c60139ad068156cef737d438cd5244742d58c2c0cb312ec61ae8dc0c463d5eb

FP8

2

NVIDIA-H200

No

Latency

c475df816df58cb8d318ea34b7bb05dff51cd2b437c2924eadf10289ab2533c8

FP8

2

NVIDIA-H200

No

Throughput

cbb3cc0300472553b8ae84bfc03f7b1f33e93ac3a9fe2dcb85783ab1981e32b1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

qwen3.5-397b-a17b#

The following table lists the supported profile configurations for qwen/qwen3.5-397b-a17b:

Precision

TP1

TP2

TP4

TP8

BF16

–

–

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

–

–

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

–

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

–

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Optimized deployment profiles#

Precision

TP

GPU

LoRA

Tuning Objective

Profile ID

NVFP4

2

NVIDIA-B200

No

Throughput

5ce5f8a3d5f0bf67afc04538745d5a9f364da7c3a79a584e84be859c911d7fb5

FP8

4

NVIDIA-H200

No

Throughput

ed6edfcd8bb730b02a1420640544dcbb13779e732f9de269b25ce1f75bcfb602

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Model-Free NIM#

vLLM#

The following models are tested and validated for nvidia/model-free-nim:

  • gpt-oss-20b

  • apriel-nemotron

  • codestral

While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.

Verified GPUs

The model-free NIM has been verified on the following GPUs:

  • NVIDIA-A100-80GB-PCIe

  • NVIDIA-A100-PCIE-40GB

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H100-PCIe

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Model-specific VRAM requirement in 2.0.12: When this NIM serves apriel-nemotron or codestral, use a GPU with at least 80 GB of usable VRAM. These models require approximately 66 GB for weights and runtime allocations and cannot start on smaller GPUs. Other validated models can use the remaining verified GPUs listed above.

SGLang#

The following models are tested and validated for nvidia/sglang-model-free-nim:

1.x NIM LLM Models#

For more information on version 1.x NIMs, refer to the 1.15 version of the NIM LLM Supported Models page.

Show 1.x models

Model (Hardware Requirements)

Organization/Model ID (Catalog Page)

Bielik 11B v2.3 Instruct

speakleash/bielik-11b-v2.3-instruct

Code Llama 13B Instruct

meta/codellama-13b-instruct

Code Llama 34B Instruct

meta/codellama-34b-instruct

Code Llama 70B Instruct

meta/codellama-70b-instruct

DeepSeek Coder V2 Lite Instruct

deepseek-ai/deepseek-coder-v2-lite-instruct

DeepSeek R1

deepseek-ai/deepseek-r1

DeepSeek R1 Distill Llama 8B

deepseek-ai/deepseek-r1-distill-llama-8b

DeepSeek R1 Distill Llama 70B

deepseek-ai/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 8B RTX

deepseek-ai/deepseek-r1-distill-llama-8b

DeepSeek R1 Distill Qwen 32B

deepseek-ai/deepseek-r1-distill-qwen-32b

DeepSeek-V3.1-Terminus

deepseek-ai/deepseek-v3.1-terminus

DeepSeek-V3.2-Exp

deepseek-ai/deepseek-v32-exp-nim

DeepSeek-V4-Flash

deepseek-ai/deepseek-v4-flash

DeepSeek-V4-Flash-0731

deepseek-ai/deepseek-v4-flash-0731

DeepSeek-V4-Pro

deepseek-ai/deepseek-v4-pro

EuroLLM 9B Instruct

utter-project/eurollm-9b-instruct

Gemma 2 2B

google/gemma-2-2b-instruct

Gemma 2 9B

google/gemma-2-9b-it

Gemma2 9B CPT Sahabat-AI v1 Instruct

gotocompany/gemma2-9b-cpt-sahabatai-v1-instruct

Gemma 3 1B Instruct

google/gemma-3-1b-it

GLM-5

zai-org/glm-5

GLM-5.1

zai-org/glm-51

GLM-5.2

zai-org/glm-5.2

GPT-OSS-20B

openai/gpt-oss-20b

GPT-OSS-120B

openai/gpt-oss-120b

Granite 3.3 8B Instruct

ibm-granite/granite-3.3-8b-instruct

GreenMind Medium 14B R1

greennode/greenmind-medium-14b-r1

Kanana 1.5 8B Instruct 2505

kakaocorp/kanana-1.5-8b-instruct-2505

(Meta) Llama 2 7B Chat

meta/llama-2-7b-chat

(Meta) Llama 2 13B Chat

meta/llama-2-13b-chat

(Meta) Llama 2 70B Chat

meta/llama-2-70b-chat

Llama 3 SQLCoder 8B

defog/llama-3-sqlcoder-8b

Llama 3 Swallow 70B Instruct V0.1

tokyotech-llm/llama-3-swallow-70b-instruct-v0.1

Llama 3 Taiwan 70B Instruct

yentinglin/llama-3-taiwan-70b-instruct

Llama 3.1 8B Base

meta/llama-3.1-8b-base

Llama 3.1 8B Instruct

meta/llama-3.1-8b-instruct

Llama-3.1-8b-Instruct-DGX-Spark

meta/llama-3.1-8b-instruct-dgx-spark

Llama-3.1-8B-Instruct PB 25h2

meta/llama-3.1-8b-instruct-pb25h2

Llama 3.1 8B Instruct RTX

meta/llama-3.1-8b-instruct

Llama 3.1 70B Instruct

meta/llama-3.1-70b-instruct

Llama-3.1-70B-Instruct PB 25h2

meta/llama-3.1-70b-instruct-pb25h2

Llama 3.1 405B Instruct

meta/llama-3.1-405b-instruct

Llama 3.1 Nemotron Nano 4B V1.1

nvidia/llama3.1-nemotron-nano-4b-v1.1

Llama 3.1 Nemotron Nano 8B V1

nvidia/llama-3.1-nemotron-nano-8b-v1

Llama-3.1-Nemotron-Nano-8B-Healthcare-Text2sql-v1.0

nvidia/llama-3.1-nemotron-nano-8b-healthcare-text2sql-v1.0

Llama 3.1 Nemotron Ultra 253B V1

nvidia/llama-3.1-nemotron-ultra-253b-v1

Llama 3.1 Nemotron 70B Instruct

nvidia/llama-3.1-nemotron-70b-instruct

Llama 3.1 Swallow 8B Instruct v0.1

tokyotech-llm/llama-3.1-swallow-8b-instruct-v0.1

Llama 3.1 Swallow 70B Instruct v0.1

tokyotech-llm/llama-3.1-swallow-70b-instruct-v0.1

Llama 3.1 Typhoon 2 8B Instruct

scb10x/llama3.1-typhoon2-8b-instruct

Llama 3.1 Typhoon 2 70B Instruct

scb10x/llama-3.1-typhoon2-70b-instruct

Llama 3.2 1B Instruct

meta/llama-3.2-1b-instruct

Llama 3.2 3B Instruct

meta/llama-3.2-3b-instruct

Llama 3.3 70B Instruct

meta/llama-3.3-70b-instruct

Llama 3.3 Nemotron Super 49B V1

nvidia/llama-3.3-nemotron-super-49b-v1

Llama-3.3-Nemotron-Super-49B-v1.5

nvidia/llama-3.3-nemotron-super-49b-v1.5

Llama-3.3-Nemotron-Super-49B-Healthcare-Text2sql-v1.0

nvidia/llama-3.3-nemotron-super-49b-healthcare-text2sql-v1.0

Llama-3.3-Nemotron-Super-49B-v1.5 PB 25h2

nvidia/llama-3.3-nemotron-super-49b-v1.5-pb25h2

Meta Llama 3 8B Instruct

meta/llama3-8b-instruct

Meta Llama 3 70B Instruct

meta/llama3-70b-instruct

Muse Glimmer

meta/muse-glimmer

MiMo-V2-Flash

xiaomi/mimo-v2-flash-experimental

MiniMax-M2.5

minimax-ai/minimax-m25

MiniMax-M2.7

minimax-ai/minimax-m27

Mistral 7B Instruct V0.3

mistralai/mistral-7b-instruct-v0.3

Mistral NeMo 12B Instruct RTX

nv-mistralai/mistral-nemo-12b-instruct

Mistral NeMo 12B Instruct

nv-mistralai/mistral-nemo-12b-instruct

Mistral NeMo Minitron 8B 8K Instruct

nv-mistralai/mistral-nemo-minitron-8b-8k-instruct

Mistral Small 24b Instruct 2501

mistralai/mistral-small-24b-instruct-2501

Mixtral 8x7B Instruct V0.1

mistralai/mixtral-8x7b-instruct-v0-1

Mixtral 8x22B Instruct V0.1

mistralai/mixtral-8x22b-instruct-v01

Nemotron 4 340B Instruct

nvidia/nemotron-4-340b-instruct

Nemotron 4 340B Reward

nvidia/nemotron-4-340b-reward

NVIDIA Nemotron 3 Nano

nvidia/nemotron-3-nano

NVIDIA-Nemotron-Nano-9B-v2

nvidia/nvidia-nemotron-nano-9b-v2

NVIDIA-Nemotron-Nano-9B-v2-DGX-Spark

nvidia/nvidia-nemotron-nano-9b-v2-dgx-spark

Nemotron-3-Super-120B-A12B

nvidia/nemotron-3-super-120b-a12b

Phi 3 Mini 4K Instruct

microsoft/phi-3-mini-4k-instruct

Phi 4 Mini Instruct

microsoft/phi-4-mini-instruct

Phind Codellama 34B V2 Instruct

phind/phind-codellama-34b-v2-instruct

Qwen3-Coder-Next

qwen/qwen3-coder-next

Qwen3-Next-80B-A3B-Instruct

qwen/qwen3-next-80b-a3b-instruct

Qwen3 Next 80B A3B Thinking

qwen/qwen3-next-80b-a3b-thinking

Qwen3-32B

qwen/qwen3-32b

Qwen3-32B NIM for DGX Spark

qwen/qwen3-32b-dgx-spark

Qwen2.5 Coder 32B Instruct

qwen/qwen2.5-coder-32b-instruct

Qwen2.5 7B Instruct

qwen/qwen-2.5-7b-instruct

Riva Translate 4B Instruct

nvidia/riva_translate_4b_instruct

Riva-Translate-4b-Instruct-v1.1

nvidia/riva-translate-4b-instruct-v1.1

Sarvam - M

sarvamai/sarvam-m

SILMA 9B Instruct v1.0

silma-ai/silma-9b-instruct-v1.0

StarCoder2 7B

bigcode/starcoder2-7b

StarCoderBase 15.5B

bigcode/starcoderbase-15b

Stockmark-2-100B-Instruct

stockmark/stockmark-2-100b-instruct

Teuken 7B Instruct Commercial v0.4

opengpt-x/teuken-7b-instruct-commercial-v0.4

Additional VLM Models#

VLMs that were released prior to version 2.0.12 are documented on the NVIDIA NIM for Vision Language Models site. Refer to that site for more information on running these VLM models.

Show additional VLM models

Model (Hardware Requirements)

Organization/Model ID (Catalog Page)

Cosmos 3 Reasoner

nvidia/cosmos3-reasoner

Cosmos Reason1 7B

nvidia/cosmos-reason1-7b

Cosmos Reason2

nvidia/cosmos-reason2-8b

DiffusionGemma 26B A4B IT

google/diffusiongemma-26b-a4b-it

Gemma 4 31B Instruct

google/gemma-4-31b-it

Gemma 4 31B IT

google/gemma-4-31b-it

Gemma-4-26B-A4B-IT

google/gemma-4-26b-a4b-it

GLM-5.3-Flash

zai-org/glm-5.3-flash

Inkling

thinkingmachines/inkling

Kimi-K2.5

moonshotai/kimi-k2.5

Kimi-K2.6 Turbo

moonshotai/kimi-k2.6-turbo

Llama 3.1 Nemotron Nano VL 8B v1

nvidia/llama-3.1-nemotron-nano-vl-8b-v1

Llama 3.2 11B Vision Instruct

meta/llama-3.2-11b-vision-instruct

Llama 3.2 90B Vision Instruct

meta/llama-3.2-90b-vision-instruct

Llama 4 Maverick 17B 128E Instruct

meta/llama-4-maverick-17b-128e-instruct

Llama 4 Scout 17B 16E Instruct

meta/llama-4-scout-17b-16e-instruct

Ministral 3 14B Instruct 2512

mistralai/ministral-3-14b-instruct-2512

Mistral Large 3 675B Instruct 2512

mistralai/mistral-large-3-675b-instruct-2512

Mistral Medium 3.5

mistralai/mistral-medium-3.5-128b

Mistral Medium 3.5 128B

mistralai/mistral-medium-3.5-128b

Mistral Small 3.2 24B Instruct 2506

mistralai/mistral-small-3.2-24b-instruct-2506

Mistral Small 4

mistralai/mistral-small-4-119b-2603

Mistral-Small-4-119B-2603

mistralai/mistral-small-4-119b-2603

Muse Glimmer

meta/muse-glimmer

nemoretriever-parse

nvidia/nemoretriever-parse

Nemotron 3.5 Content Safety

nvidia/nemotron-3.5-content-safety

Nemotron Nano 12B v2 VL

nvidia/nemotron-nano-12b-v2-vl

Nemotron Parse

nvidia/nemotron-parse

Nemotron Parse v2.0

nvidia/nemotron-parse-v2.0

Nemotron-3-Content-Safety

nvidia/nemotron-3-content-safety

Nemotron-3-Nano-Omni-30B-A3B-Reasoning

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning

Nemotron-Parse-v1.2

nvidia/nemotron-parse-v1.2

NVIDIA Nemotron 3 Nano Omni

nvidia/nemotron-3-nano-omni-30b-a3b-reasoning

Qwen3.5

qwen/qwen3.5-35b-a3b

Qwen3.5-122B-A10B

qwen/qwen3.5-122b-a10b

Qwen3.5-397B-A17B

qwen/qwen3.5-397b-a17b

Qwen3.6-27B

qwen/qwen3.6-27b

Qwen3.6-35B-A3B

qwen/qwen3.6-35b-a3b

Qwen3.8-27B

qwen/qwen3.8-27b

Qwen3.8-Flash-Next

nvidia/vllm-model-free-nim

Step 3.7 Flash

stepfun-ai/step-3.7-flash