Support Matrix for NIMs#

This page lists supported models, deployment profiles, and verified hardware SKUs for NIM LLM and VLM. This page is organized according to model type:

Use the following sections to identify the supported deployment profiles for each model. Profile strings follow a naming convention described in Model Profiles and Selection.

Important

The list of NIM Certified models is available on the NGC Catalog site.

Use the table below to filter NIM profiles by GPU, tensor parallelism (TP), precision, model type, LoRA support, and model name. Each row is one supported profile; details also appear in the per-model sections that follow.

Note

The GB300-WS GPU is more commonly known as DGX Station.

ModelTPPrecisionLoRA
deepseek-v4-pro-0813 8 FP8 No
gpt-oss-120b 1 MXFP4 No
gpt-oss-120b 2 MXFP4 No
gpt-oss-120b 4 MXFP4 No
gpt-oss-120b 8 MXFP4 No
gpt-oss-120b 1 MXFP4 Yes
gpt-oss-120b 2 MXFP4 Yes
gpt-oss-120b 4 MXFP4 Yes
gpt-oss-120b 8 MXFP4 Yes
gpt-oss-20b 1 MXFP4 No
gpt-oss-20b 2 MXFP4 No
gpt-oss-20b 4 MXFP4 No
gpt-oss-20b 8 MXFP4 No
gpt-oss-20b 1 MXFP4 Yes
gpt-oss-20b 2 MXFP4 Yes
gpt-oss-20b 4 MXFP4 Yes
gpt-oss-20b 8 MXFP4 Yes
kimi-k2.6 4 NVFP4 No
kimi-k2.6 8 INT4 No
llama-3.1-70b-instruct 1 BF16 No
llama-3.1-70b-instruct 2 BF16 No
llama-3.1-70b-instruct 4 BF16 No
llama-3.1-70b-instruct 8 BF16 No
llama-3.1-70b-instruct 1 BF16 Yes
llama-3.1-70b-instruct 2 BF16 Yes
llama-3.1-70b-instruct 4 BF16 Yes
llama-3.1-70b-instruct 8 BF16 Yes
llama-3.1-70b-instruct 1 FP8 No
llama-3.1-70b-instruct 2 FP8 No
llama-3.1-70b-instruct 4 FP8 No
llama-3.1-70b-instruct 8 FP8 No
llama-3.1-70b-instruct 1 FP8 Yes
llama-3.1-70b-instruct 2 FP8 Yes
llama-3.1-70b-instruct 4 FP8 Yes
llama-3.1-70b-instruct 8 FP8 Yes
llama-3.1-70b-instruct 1 NVFP4 No
llama-3.1-70b-instruct 2 NVFP4 No
llama-3.1-70b-instruct 4 NVFP4 No
llama-3.1-70b-instruct 8 NVFP4 No
llama-3.1-70b-instruct 1 NVFP4 Yes
llama-3.1-70b-instruct 2 NVFP4 Yes
llama-3.1-70b-instruct 4 NVFP4 Yes
llama-3.1-70b-instruct 8 NVFP4 Yes
llama-3.1-8b-instruct 1 BF16 No
llama-3.1-8b-instruct 1 BF16 Yes
llama-3.1-8b-instruct 1 FP8 No
llama-3.1-8b-instruct 1 FP8 Yes
llama-3.1-8b-instruct 1 NVFP4 No
llama-3.1-8b-instruct 1 NVFP4 Yes
llama-3.3-70b-instruct 1 BF16 No
llama-3.3-70b-instruct 2 BF16 No
llama-3.3-70b-instruct 4 BF16 No
llama-3.3-70b-instruct 8 BF16 No
llama-3.3-70b-instruct 1 BF16 Yes
llama-3.3-70b-instruct 2 BF16 Yes
llama-3.3-70b-instruct 4 BF16 Yes
llama-3.3-70b-instruct 8 BF16 Yes
llama-3.3-70b-instruct 1 FP8 No
llama-3.3-70b-instruct 2 FP8 No
llama-3.3-70b-instruct 4 FP8 No
llama-3.3-70b-instruct 8 FP8 No
llama-3.3-70b-instruct 1 FP8 Yes
llama-3.3-70b-instruct 2 FP8 Yes
llama-3.3-70b-instruct 4 FP8 Yes
llama-3.3-70b-instruct 8 FP8 Yes
llama-3.3-70b-instruct 1 NVFP4 No
llama-3.3-70b-instruct 2 NVFP4 No
llama-3.3-70b-instruct 4 NVFP4 No
llama-3.3-70b-instruct 8 NVFP4 No
llama-3.3-70b-instruct 1 NVFP4 Yes
llama-3.3-70b-instruct 2 NVFP4 Yes
llama-3.3-70b-instruct 4 NVFP4 Yes
llama-3.3-70b-instruct 8 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 1 BF16 No
llama-3.3-nemotron-super-49b-v1.5 2 BF16 No
llama-3.3-nemotron-super-49b-v1.5 4 BF16 No
llama-3.3-nemotron-super-49b-v1.5 8 BF16 No
llama-3.3-nemotron-super-49b-v1.5 1 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 2 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 4 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 8 BF16 Yes
llama-3.3-nemotron-super-49b-v1.5 1 FP8 No
llama-3.3-nemotron-super-49b-v1.5 2 FP8 No
llama-3.3-nemotron-super-49b-v1.5 4 FP8 No
llama-3.3-nemotron-super-49b-v1.5 8 FP8 No
llama-3.3-nemotron-super-49b-v1.5 1 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 2 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 4 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 8 FP8 Yes
llama-3.3-nemotron-super-49b-v1.5 1 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 2 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 4 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 8 NVFP4 No
llama-3.3-nemotron-super-49b-v1.5 1 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 2 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 4 NVFP4 Yes
llama-3.3-nemotron-super-49b-v1.5 8 NVFP4 Yes
mistral-small-4-119b-2603 1 FP8 No
mistral-small-4-119b-2603 2 FP8 No
mistral-small-4-119b-2603 4 FP8 No
mistral-small-4-119b-2603 8 FP8 No
mistral-small-4-119b-2603 1 NVFP4 No
mistral-small-4-119b-2603 2 NVFP4 No
mistral-small-4-119b-2603 4 NVFP4 No
mistral-small-4-119b-2603 8 NVFP4 No
nemotron-content-safety-reasoning-4b-experimental 1 BF16 No
nemotron-content-safety-reasoning-4b-experimental 2 BF16 No
nemotron-content-safety-reasoning-4b-experimental 4 BF16 No
nemotron-content-safety-reasoning-4b-experimental 8 BF16 No
nemotron-3-nano 1 BF16 No
nemotron-3-nano 2 BF16 No
nemotron-3-nano 4 BF16 No
nemotron-3-nano 8 BF16 No
nemotron-3-nano 1 BF16 Yes
nemotron-3-nano 2 BF16 Yes
nemotron-3-nano 4 BF16 Yes
nemotron-3-nano 8 BF16 Yes
nemotron-3-nano 1 FP8 No
nemotron-3-nano 2 FP8 No
nemotron-3-nano 4 FP8 No
nemotron-3-nano 8 FP8 No
nemotron-3-nano 1 NVFP4 No
nemotron-3-nano 2 NVFP4 No
nemotron-3-nano 4 NVFP4 No
nemotron-3-nano 8 NVFP4 No
nemotron-3-nano 1 NVFP4 Yes
nemotron-3-nano-omni-30b-a3b-reasoning 2 BF16 No
nemotron-3-nano-omni-30b-a3b-reasoning 1 FP8 No
nemotron-3-nano-omni-30b-a3b-reasoning 1 NVFP4 No
nemotron-3-super-120b-a12b 1 BF16 No
nemotron-3-super-120b-a12b 2 BF16 No
nemotron-3-super-120b-a12b 4 BF16 No
nemotron-3-super-120b-a12b 8 BF16 No
nemotron-3-super-120b-a12b 2 BF16 Yes
nemotron-3-super-120b-a12b 4 BF16 Yes
nemotron-3-super-120b-a12b 8 BF16 Yes
nemotron-3-super-120b-a12b 1 FP8 No
nemotron-3-super-120b-a12b 2 FP8 No
nemotron-3-super-120b-a12b 4 FP8 No
nemotron-3-super-120b-a12b 8 FP8 No
nemotron-3-super-120b-a12b 1 NVFP4 No
nemotron-3-super-120b-a12b 2 NVFP4 No
nemotron-3-super-120b-a12b 4 NVFP4 No
nemotron-3-super-120b-a12b 8 NVFP4 No
nemotron-3-super-120b-a12b 1 NVFP4 Yes
nemotron-3-super-120b-a12b 2 NVFP4 Yes
nemotron-3-ultra-550b-a55b 8 BF16 No
nemotron-3-ultra-550b-a55b 2 NVFP4 No
nemotron-3-ultra-550b-a55b 4 NVFP4 No
nemotron-3-ultra-550b-a55b 8 NVFP4 No
nemotron-3.5-lightning-30b-a3b 1 BF16 No
nemotron-3.5-lightning-30b-a3b 1 BF16 Yes
nemotron-3.5-lightning-30b-a3b 2 BF16 No
nemotron-3.5-lightning-30b-a3b 2 BF16 Yes
nemotron-3.5-lightning-30b-a3b 4 BF16 No
nemotron-3.5-lightning-30b-a3b 4 BF16 Yes
nemotron-3.5-lightning-30b-a3b 8 BF16 No
nemotron-3.5-lightning-30b-a3b 8 BF16 Yes
nemotron-3.5-lightning-30b-a3b 1 W4A16 No
nemotron-3.5-lightning-30b-a3b 1 W4A16 Yes
nemotron-3.5-lightning-30b-a3b 2 W4A16 No
nemotron-3.5-lightning-30b-a3b 2 W4A16 Yes
nemotron-3.5-lightning-30b-a3b 4 W4A16 No
nemotron-3.5-lightning-30b-a3b 4 W4A16 Yes
nemotron-3.5-lightning-30b-a3b 1 NVFP4 No
nemotron-3.5-lightning-30b-a3b 1 NVFP4 Yes
nemotron-3.5-lightning-30b-a3b 2 NVFP4 No
nemotron-3.5-lightning-30b-a3b 2 NVFP4 Yes
nemotron-3.5-lightning-30b-a3b 4 NVFP4 No
nemotron-3.5-lightning-30b-a3b 4 NVFP4 Yes
qwen3.5-122b-a10b 1 BF16 No
qwen3.5-122b-a10b 1 BF16 Yes
qwen3.5-122b-a10b 2 BF16 No
qwen3.5-122b-a10b 2 BF16 Yes
qwen3.5-122b-a10b 4 BF16 No
qwen3.5-122b-a10b 4 BF16 Yes
qwen3.5-122b-a10b 8 BF16 No
qwen3.5-122b-a10b 8 BF16 Yes
qwen3.5-122b-a10b 1 FP8 No
qwen3.5-122b-a10b 1 FP8 Yes
qwen3.5-122b-a10b 2 FP8 No
qwen3.5-122b-a10b 2 FP8 Yes
qwen3.5-122b-a10b 4 FP8 No
qwen3.5-122b-a10b 4 FP8 Yes
qwen3.5-122b-a10b 8 FP8 No
qwen3.5-122b-a10b 8 FP8 Yes
qwen3.5-122b-a10b 1 NVFP4 No
qwen3.5-122b-a10b 1 NVFP4 Yes
qwen3.5-122b-a10b 2 NVFP4 No
qwen3.5-122b-a10b 2 NVFP4 Yes
qwen3.5-122b-a10b 4 NVFP4 No
qwen3.5-122b-a10b 4 NVFP4 Yes
qwen3.5-122b-a10b 8 NVFP4 No
qwen3.5-122b-a10b 8 NVFP4 Yes
qwen3.5-397b-a17b 4 BF16 No
qwen3.5-397b-a17b 4 BF16 Yes
qwen3.5-397b-a17b 8 BF16 No
qwen3.5-397b-a17b 8 BF16 Yes
qwen3.5-397b-a17b 2 FP8 No
qwen3.5-397b-a17b 2 FP8 Yes
qwen3.5-397b-a17b 4 FP8 No
qwen3.5-397b-a17b 4 FP8 Yes
qwen3.5-397b-a17b 8 FP8 No
qwen3.5-397b-a17b 8 FP8 Yes
qwen3.5-397b-a17b 1 NVFP4 No
qwen3.5-397b-a17b 1 NVFP4 Yes
qwen3.5-397b-a17b 2 NVFP4 No
qwen3.5-397b-a17b 2 NVFP4 Yes
qwen3.5-397b-a17b 4 NVFP4 No
qwen3.5-397b-a17b 4 NVFP4 Yes
qwen3.5-397b-a17b 8 NVFP4 No
qwen3.5-397b-a17b 8 NVFP4 Yes
riva-translate-4b-instruct-v2 1 BF16 No
starcoder2-7b 1 BF16 No
starcoder2-7b 2 BF16 No

LLMs#

The following section lists all of the supported LLMs.

deepseek-v4-pro-0813#

The following table lists the supported profile configurations for deepseek-ai/DeepSeek-V4-Pro-0813.

This NIM is an SGLang 2.1.2-variant container. Combined container and model disk space is approximately 831 GB. The fallback profile requires more than 1,024 GB of aggregate GPU memory across eight GPUs. For more information, refer to the dedicated Get Started guide.

GPU

GPU Memory

Precision

TP

PP

LoRA

Profile

Any (fallback)

> 1,024 GB aggregate

FP8

8

1

No

sglang-fp8-tp8-fallback-1024g

H20-3e

141 GB

FP8

8

1

No

sglang-h20-3e-fp8-tp8-throughput-232c:10de-1024g

H200

141 GB

FP8

8

1

No

sglang-h200-fp8-tp8-throughput-2335:10de-1024g

B200

196 GB

FP8

8

1

No

sglang-b200-fp8-tp8-throughput-2901:10de-1024g

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-H20-3e

  • NVIDIA-H200

gpt-oss-120b#

The following table lists the supported profile configurations for openai/gpt-oss-120b:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Support requirements in 2.0.12: The MXFP4 TP8 profile on NVIDIA A10G and NVIDIA A100-SXM4-40GB requires NIM_KVCACHE_PERCENT=0.80. The LoRA TP4 and TP8 profiles on NVIDIA L40S and NVIDIA RTX PRO 4500 Blackwell Server Edition require piecewise CUDA graphs. LoRA adapters with rank greater than 128 are not supported.

gpt-oss-20b#

The following table lists the supported profile configurations for openai/gpt-oss-20b:

Precision

TP1

TP2

TP4

TP8

MXFP4

vllm-mxfp4-tp1-pp1

vllm-mxfp4-tp2-pp1

vllm-mxfp4-tp4-pp1

vllm-mxfp4-tp8-pp1

MXFP4 + LoRA

vllm-mxfp4-tp1-pp1-lora

vllm-mxfp4-tp2-pp1-lora

vllm-mxfp4-tp4-pp1-lora

vllm-mxfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Support requirements in 2.0.12: LoRA adapters with rank greater than 128 are not supported. On NVIDIA A10G, LoRA deployments require NIM_MAX_MODEL_LEN below the default 131,072 token context length so that the KV cache fits in device memory.

llama-3.1-70b-instruct#

The following table lists the supported profile configurations for meta/llama-3.1-70b-instruct:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Support change in 2.0.11, still applies in 2.0.12: BF16 TP2 with or without LoRA is not supported on NVIDIA RTX PRO 6000 Blackwell Server Edition, and BF16 TP8 without LoRA is not supported on NVIDIA B200. These combinations are omitted from the profile-to-GPU map above.

llama-3.1-8b-instruct#

The following table lists the supported profile configurations for meta/llama-3.1-8b-instruct:

Precision

TP1

BF16

vllm-bf16-tp1-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

FP8

vllm-fp8-tp1-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-70b-instruct#

The following table lists the supported profile configurations for meta/llama-3.3-70b-instruct:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

llama-3.3-nemotron-super-49b-v1.5#

The following table lists the supported profile configurations for nvidia/llama-3.3-nemotron-super-49b-v1.5:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

nemotron-content-safety-reasoning-4b-experimental#

The following table lists the supported profile configurations for nvidia/nemotron-content-safety-reasoning-4b-experimental. The latest release version of this NIM is 2.0.3. For more information, refer to the dedicated Get Started guide.

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-GB200

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

nemotron-3-nano#

The following table lists the supported profile configurations for nvidia/nemotron-3-nano:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

–

–

–

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Support requirement in 2.0.12: The fused mixture-of-experts kernels support LoRA adapters with a maximum rank of 128. Adapters with a greater rank are not supported.

nemotron-3-super-120b-a12b#

The following table lists the supported profile configurations for nvidia/nemotron-3-super-120b-a12b:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

–

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

–

–

Note

This is a large model. Lower-TP profiles require substantially more GPU memory per device, so some verified GPUs support only TP4 or TP8 profiles.

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Support change in 2.0.11, still applies in 2.0.12: The following BF16 LoRA combinations are not supported: TP2 on NVIDIA GH200-144G-HBM3e and NVIDIA H200-NVL, TP4 on NVIDIA H100-NVL, and TP8 on NVIDIA L40S. These combinations are omitted from the profile-to-GPU map above.

nemotron-3-ultra-550b-a55b#

The following table lists the supported profile configurations for nvidia/nemotron-3-ultra-550b-a55b:

Precision

TP1

TP2

TP4

TP8

BF16

–

–

–

vllm-bf16-tp8-pp1

BF16 + LoRA

–

–

–

–

NVFP4

–

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

Note

The minimum deployment is two GPUs, using vllm-nvfp4-tp2-pp1 on high-memory Blackwell SKUs (NVIDIA-B300-SXM6-AC and NVIDIA-GB300); the other profiles use four or eight GPUs. H200 validation covers one NVFP4 TP4 profile and is less extensive than the Blackwell validation.

FP8 profiles were not released for this NIM because NVFP4 profiles work on all tested Hopper-architecture GPUs (for example, H100 and H200) and deliver better performance than FP8 profiles.

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

nemotron-3.5-lightning-30b-a3b#

The following table lists the supported profile configurations for nvidia/nemotron-3.5-lightning-30b-a3b, along with the minimum GPU VRAM required per device and the GPU architecture that each precision requires. The VRAM values come from the profile disk-sizing metadata reported by list-model-profiles and represent the minimum per-GPU memory needed to load the model weights and the runtime allocations; they do not include additional headroom for large KV caches at high context length or high concurrency.

The latest release version of this NIM is 2.0.10. For more information, refer to the dedicated Get Started guide.

Nemotron 3.5 Lightning Profile Information#

Precision

TP

PP

Min VRAM per GPU

Min GPU count

Architecture requirement

LoRA Profile

BF16

1

1

66 GB

1

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

2

1

35 GB

2

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

4

1

20 GB

4

Ampere or newer (SM 8.0+)

Base and LoRA

BF16

8

1

12 GB

8

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

1

1

32 GB

1

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

2

1

20 GB

2

Ampere or newer (SM 8.0+)

Base and LoRA

W4A16

4

1

14 GB

4

Ampere or newer (SM 8.0+)

Base and LoRA

NVFP4

1

1

30 GB

1

Blackwell or newer (SM 10.0+)

Base and LoRA

NVFP4

2

1

18 GB

2

Blackwell or newer (SM 10.0+)

Base and LoRA

NVFP4

4

1

12 GB

4

Blackwell or newer (SM 10.0+)

Base and LoRA

Notes:

  • Min VRAM per GPU is the floor that list-model-profiles reports for that precision and tensor-parallel size. A GPU below the floor is filtered out during profile selection.

  • Architecture requirement reflects the kernels each precision needs. NVFP4 requires Streaming Multiprocessor (SM) 10.0+ (Blackwell) because it uses NVFP4-native tensor cores. BF16 and W4A16 run on any Ampere-class or newer GPU (SM 8.0+).

  • Sharing weights across more GPUs (higher TP) lowers the per-GPU VRAM floor and enables lower-memory GPUs.

  • Model-cache disk sizes range from approximately 19 GB (NVFP4) to approximately 63 GB (BF16). Cache size is separate from GPU VRAM and does not indicate that the model will fit in a given GPU.

Verified GPUs

This model has been verified on the following GPUs (grouped by architecture):

Blackwell (SM 10.0) — supports NVFP4, W4A16, BF16:

GPU

VRAM per GPU

Profiles that fit

NVIDIA-B200

180 GB

All profiles

NVIDIA-B300-SXM6-AC

288 GB

All profiles

NVIDIA-GB200

186 GB

All profiles

NVIDIA-GB300

288 GB

All profiles

NVIDIA-GB10

128 GB (unified memory)

All profiles; use --gpu-memory-utilization 0.75 on GB10 (see the getting started guide)

NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

96 GB

All profiles

Hopper (SM 9.0) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-H100-80GB-HBM3

80 GB

BF16 all sizes (needs --max-num-seqs 512 on TP=1), W4A16 all sizes

NVIDIA-H200

141 GB

BF16 all sizes, W4A16 all sizes

Ampere (SM 8.0) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-A100-SXM4-80GB

80 GB

BF16 TP≥2, W4A16 all sizes

Ada (SM 8.9) — supports W4A16 and BF16 (no NVFP4):

GPU

VRAM per GPU

Profiles that fit

NVIDIA-L40S

48 GB

BF16 TP≥2, W4A16 all sizes

Unverified but potentially compatible. Any GPU that meets the per-profile minimum VRAM floor and the architecture requirement in the profile table can run that profile technically, even if it is not in the preceding Verified GPUs tables. Only the GPUs in those tables have been tested end-to-end for this model at this NIM version; untested systems may exhibit different startup, throughput, or accuracy behavior.

riva-translate-4b-instruct-v2#

The following table lists the supported profile configuration for nvidia/riva-translate-4b-instruct-v2.

The latest release version of this NIM is 2.0.8.

Precision

TP

PP

LoRA

Profile

BF16

1

1

No

vllm-bf16-tp1-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

starcoder2-7b#

The following table lists the supported profile configurations for bigcode/starcoder2-7b:

Precision

TP1

TP2

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-80GB-PCIe

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-GB300-WS

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

VLMs#

The following section lists all of the supported VLMs.

kimi-k2.6#

The following table lists the supported profile configurations for moonshotai/kimi-k2.6:

Precision

TP1

TP2

TP4

TP8

NVFP4

–

–

vllm-nvfp4-tp4-pp1

–

INT4

–

–

–

vllm-int4-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H200

mistral-small-4-119b-2603#

The following table lists the supported profile configurations for mistralai/mistral-small-4-119b-2603:

Precision

TP1

TP2

TP4

TP8

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H200-NVL

nemotron-3-nano-omni-30b-a3b-reasoning#

The following table lists the supported profile configurations for nvidia/nemotron-3-nano-omni-30b-a3b-reasoning:

Precision

TP1

TP2

TP4

TP8

BF16

–

vllm-bf16-tp2-pp1

–

–

FP8

vllm-fp8-tp1-pp1

–

–

–

NVFP4

vllm-nvfp4-tp1-pp1

–

–

–

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

qwen3.5-122b-a10b#

The following table lists the supported profile configurations for qwen/qwen3.5-122b-a10b:

Precision

TP1

TP2

TP4

TP8

BF16

vllm-bf16-tp1-pp1

vllm-bf16-tp2-pp1

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

vllm-bf16-tp1-pp1-lora

vllm-bf16-tp2-pp1-lora

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

vllm-fp8-tp1-pp1

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

vllm-fp8-tp1-pp1-lora

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

qwen3.5-397b-a17b#

The following table lists the supported profile configurations for qwen/qwen3.5-397b-a17b:

Precision

TP1

TP2

TP4

TP8

BF16

–

–

vllm-bf16-tp4-pp1

vllm-bf16-tp8-pp1

BF16 + LoRA

–

–

vllm-bf16-tp4-pp1-lora

vllm-bf16-tp8-pp1-lora

FP8

–

vllm-fp8-tp2-pp1

vllm-fp8-tp4-pp1

vllm-fp8-tp8-pp1

FP8 + LoRA

–

vllm-fp8-tp2-pp1-lora

vllm-fp8-tp4-pp1-lora

vllm-fp8-tp8-pp1-lora

NVFP4

vllm-nvfp4-tp1-pp1

vllm-nvfp4-tp2-pp1

vllm-nvfp4-tp4-pp1

vllm-nvfp4-tp8-pp1

NVFP4 + LoRA

vllm-nvfp4-tp1-pp1-lora

vllm-nvfp4-tp2-pp1-lora

vllm-nvfp4-tp4-pp1-lora

vllm-nvfp4-tp8-pp1-lora

Verified GPUs

This model has been verified on the following GPUs:

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Model-Free NIM#

vLLM#

The following models are tested and validated for nvidia/model-free-nim:

  • gpt-oss-20b

  • apriel-nemotron

  • codestral

While not explicitly validated, the model-free NIM can be used with any model supported by the underlying backend (vLLM) version. Refer to Model-Free NIM for deployment details.

Verified GPUs

The model-free NIM has been verified on the following GPUs:

  • NVIDIA-A100-80GB-PCIe

  • NVIDIA-A100-PCIE-40GB

  • NVIDIA-A100-SXM4-40GB

  • NVIDIA-A100-SXM4-80GB

  • NVIDIA-A10G

  • NVIDIA-B200

  • NVIDIA-B300-SXM6-AC

  • NVIDIA-GB10

  • NVIDIA-GB200

  • NVIDIA-GB300

  • NVIDIA-GB300-WS

  • NVIDIA-GH200-144G-HBM3e

  • NVIDIA-GH200-480GB

  • NVIDIA-H100-80GB-HBM3

  • NVIDIA-H100-NVL

  • NVIDIA-H100-PCIe

  • NVIDIA-H200

  • NVIDIA-H200-NVL

  • NVIDIA-L40S

  • NVIDIA-RTX-PRO-4500-Blackwell-Server-Edition

  • NVIDIA-RTX-PRO-6000-Blackwell-Server-Edition

Note

Model-specific VRAM requirement in 2.0.12: When this NIM serves apriel-nemotron or codestral, use a GPU with at least 80 GB of usable VRAM. These models require approximately 66 GB for weights and runtime allocations and cannot start on smaller GPUs. Other validated models can use the remaining verified GPUs listed above.

SGLang#

The following models are tested and validated for nvidia/sglang-model-free-nim:

1.x NIM LLM and VLM Models#

For more information on version 1.x NIMs, refer to the 1.15 version of the NIM LLM and VLM Supported Models page.

Show 1.x models

Model (Hardware Requirements)

Organization/Model ID (Catalog Page)

Bielik 11B v2.3 Instruct

speakleash/bielik-11b-v2.3-instruct

Code Llama 13B Instruct

meta/codellama-13b-instruct

Code Llama 34B Instruct

meta/codellama-34b-instruct

Code Llama 70B Instruct

meta/codellama-70b-instruct

DeepSeek Coder V2 Lite Instruct

deepseek-ai/deepseek-coder-v2-lite-instruct

DeepSeek R1

deepseek-ai/deepseek-r1

DeepSeek R1 Distill Llama 8B

deepseek-ai/deepseek-r1-distill-llama-8b

DeepSeek R1 Distill Llama 70B

deepseek-ai/deepseek-r1-distill-llama-70b

DeepSeek R1 Distill Llama 8B RTX

deepseek-ai/deepseek-r1-distill-llama-8b

DeepSeek R1 Distill Qwen 32B

deepseek-ai/deepseek-r1-distill-qwen-32b

DeepSeek-V3.1-Terminus

deepseek-ai/deepseek-v3.1-terminus

DeepSeek-V3.2-Exp

deepseek-ai/deepseek-v32-exp-nim

DeepSeek-V4-Flash

deepseek-ai/deepseek-v4-flash

DeepSeek-V4-Flash-0731

deepseek-ai/deepseek-v4-flash-0731

DeepSeek-V4-Pro

deepseek-ai/deepseek-v4-pro

EuroLLM 9B Instruct

utter-project/eurollm-9b-instruct

Gemma 2 2B

google/gemma-2-2b-instruct

Gemma 2 9B

google/gemma-2-9b-it

Gemma2 9B CPT Sahabat-AI v1 Instruct

gotocompany/gemma2-9b-cpt-sahabatai-v1-instruct

Gemma 3 1B Instruct

google/gemma-3-1b-it

GLM-5

zai-org/glm-5

GLM-5.1

zai-org/glm-51

GLM-5.2

zai-org/glm-5.2

GPT-OSS-20B

openai/gpt-oss-20b

GPT-OSS-120B

openai/gpt-oss-120b

Granite 3.3 8B Instruct

ibm-granite/granite-3.3-8b-instruct

GreenMind Medium 14B R1

greennode/greenmind-medium-14b-r1

Kanana 1.5 8B Instruct 2505

kakaocorp/kanana-1.5-8b-instruct-2505

(Meta) Llama 2 7B Chat

meta/llama-2-7b-chat

(Meta) Llama 2 13B Chat

meta/llama-2-13b-chat

(Meta) Llama 2 70B Chat

meta/llama-2-70b-chat

Llama 3 SQLCoder 8B

defog/llama-3-sqlcoder-8b

Llama 3 Swallow 70B Instruct V0.1

tokyotech-llm/llama-3-swallow-70b-instruct-v0.1

Llama 3 Taiwan 70B Instruct

yentinglin/llama-3-taiwan-70b-instruct

Llama 3.1 8B Base

meta/llama-3.1-8b-base

Llama 3.1 8B Instruct

meta/llama-3.1-8b-instruct

Llama-3.1-8b-Instruct-DGX-Spark

meta/llama-3.1-8b-instruct-dgx-spark

Llama-3.1-8B-Instruct PB 25h2

meta/llama-3.1-8b-instruct-pb25h2

Llama 3.1 8B Instruct RTX

meta/llama-3.1-8b-instruct

Llama 3.1 70B Instruct

meta/llama-3.1-70b-instruct

Llama-3.1-70B-Instruct PB 25h2

meta/llama-3.1-70b-instruct-pb25h2

Llama 3.1 405B Instruct

meta/llama-3.1-405b-instruct

Llama 3.1 Nemotron Nano 4B V1.1

nvidia/llama3.1-nemotron-nano-4b-v1.1

Llama 3.1 Nemotron Nano 8B V1

nvidia/llama-3.1-nemotron-nano-8b-v1

Llama-3.1-Nemotron-Nano-8B-Healthcare-Text2sql-v1.0

nvidia/llama-3.1-nemotron-nano-8b-healthcare-text2sql-v1.0

Llama 3.1 Nemotron Ultra 253B V1

nvidia/llama-3.1-nemotron-ultra-253b-v1

Llama 3.1 Nemotron 70B Instruct

nvidia/llama-3.1-nemotron-70b-instruct

Llama 3.1 Swallow 8B Instruct v0.1

tokyotech-llm/llama-3.1-swallow-8b-instruct-v0.1

Llama 3.1 Swallow 70B Instruct v0.1

tokyotech-llm/llama-3.1-swallow-70b-instruct-v0.1

Llama 3.1 Typhoon 2 8B Instruct

scb10x/llama3.1-typhoon2-8b-instruct

Llama 3.1 Typhoon 2 70B Instruct

scb10x/llama-3.1-typhoon2-70b-instruct

Llama 3.2 1B Instruct

meta/llama-3.2-1b-instruct

Llama 3.2 3B Instruct

meta/llama-3.2-3b-instruct

Llama 3.3 70B Instruct

meta/llama-3.3-70b-instruct

Llama 3.3 Nemotron Super 49B V1

nvidia/llama-3.3-nemotron-super-49b-v1

Llama-3.3-Nemotron-Super-49B-v1.5

nvidia/llama-3.3-nemotron-super-49b-v1.5

Llama-3.3-Nemotron-Super-49B-Healthcare-Text2sql-v1.0

nvidia/llama-3.3-nemotron-super-49b-healthcare-text2sql-v1.0

Llama-3.3-Nemotron-Super-49B-v1.5 PB 25h2

nvidia/llama-3.3-nemotron-super-49b-v1.5-pb25h2

Meta Llama 3 8B Instruct

meta/llama3-8b-instruct

Meta Llama 3 70B Instruct

meta/llama3-70b-instruct

Muse Glimmer

meta/muse-glimmer

MiMo-V2-Flash

xiaomi/mimo-v2-flash-experimental

MiniMax-M2.5

minimax-ai/minimax-m25

MiniMax-M2.7

minimax-ai/minimax-m27

Mistral 7B Instruct V0.3

mistralai/mistral-7b-instruct-v0.3

Mistral NeMo 12B Instruct RTX

nv-mistralai/mistral-nemo-12b-instruct

Mistral NeMo 12B Instruct

nv-mistralai/mistral-nemo-12b-instruct

Mistral NeMo Minitron 8B 8K Instruct

nv-mistralai/mistral-nemo-minitron-8b-8k-instruct

Mistral Small 24b Instruct 2501

mistralai/mistral-small-24b-instruct-2501

Mixtral 8x7B Instruct V0.1

mistralai/mixtral-8x7b-instruct-v0-1

Mixtral 8x22B Instruct V0.1

mistralai/mixtral-8x22b-instruct-v01

Nemotron 4 340B Instruct

nvidia/nemotron-4-340b-instruct

Nemotron 4 340B Reward

nvidia/nemotron-4-340b-reward

NVIDIA Nemotron 3 Nano

nvidia/nemotron-3-nano

NVIDIA-Nemotron-Nano-9B-v2

nvidia/nvidia-nemotron-nano-9b-v2

NVIDIA-Nemotron-Nano-9B-v2-DGX-Spark

nvidia/nvidia-nemotron-nano-9b-v2-dgx-spark

Nemotron-3-Super-120B-A12B

nvidia/nemotron-3-super-120b-a12b

Phi 3 Mini 4K Instruct

microsoft/phi-3-mini-4k-instruct

Phi 4 Mini Instruct

microsoft/phi-4-mini-instruct

Phind Codellama 34B V2 Instruct

phind/phind-codellama-34b-v2-instruct

Qwen3-Coder-Next

qwen/qwen3-coder-next

Qwen3-Next-80B-A3B-Instruct

qwen/qwen3-next-80b-a3b-instruct

Qwen3 Next 80B A3B Thinking

qwen/qwen3-next-80b-a3b-thinking

Qwen3-32B

qwen/qwen3-32b

Qwen3-32B NIM for DGX Spark

qwen/qwen3-32b-dgx-spark

Qwen2.5 Coder 32B Instruct

qwen/qwen2.5-coder-32b-instruct

Qwen2.5 7B Instruct

qwen/qwen-2.5-7b-instruct

Riva Translate 4B Instruct

nvidia/riva_translate_4b_instruct

Riva-Translate-4b-Instruct-v1.1

nvidia/riva-translate-4b-instruct-v1.1

Sarvam - M

sarvamai/sarvam-m

SILMA 9B Instruct v1.0

silma-ai/silma-9b-instruct-v1.0

StarCoder2 7B

bigcode/starcoder2-7b

StarCoderBase 15.5B

bigcode/starcoderbase-15b

Stockmark-2-100B-Instruct

stockmark/stockmark-2-100b-instruct

Teuken 7B Instruct Commercial v0.4

opengpt-x/teuken-7b-instruct-commercial-v0.4