Support Matrix#
Hardware#
Unless specified otherwise, NVIDIA NIM for vision language models (VLMs) should, but are not guaranteed to, run on any NVIDIA GPU, provided the GPU has sufficient memory. They can also run on multiple homogeneous NVIDIA GPUs with sufficient aggregate memory and a CUDA compute capability of >= 7.0 (8.0 for bfloat16) unless otherwise specified. For more information, refer to Supported Models.
NVIDIA NIM for VLMs does not support NVIDIA Virtual GPU (vGPU) environments.
For information on the supported operating systems, drivers, and software, refer to the About Get Started page.
Supported Models#
Qwen3.5-397B-A17B#
Latest supported release tag: 2.0.4-variant
The following section lists the supported configurations for
qwen/qwen3.5-397b-a17b (NGC catalog page).
Generic Configuration#
NIM for VLMs offers competitive performance through a custom vLLM backend. Any NVIDIA GPU with sufficient memory should be able to run this model, though this is not guaranteed.
The GPU Memory and Disk Space values are in GB. The disk space values for LoRA-enabled profiles do not include the LoRA model weights.
GPU |
GPU Memory |
Precision |
LoRA |
# of GPUs |
Disk Space |
|---|---|---|---|---|---|
B300 |
288 |
NVFP4 |
Yes |
2,4,8 |
240 |
B300 |
288 |
FP8 |
Yes |
4,8 |
380 |
B300 |
288 |
BF16 |
Yes |
4,8 |
760 |
B300 |
288 |
NVFP4 |
No |
2,4,8 |
240 |
B300 |
288 |
FP8 |
No |
2,4,8 |
380 |
B300 |
288 |
BF16 |
No |
4,8 |
760 |
GB300 |
288 |
NVFP4 |
Yes |
2,4 |
240 |
GB300 |
288 |
FP8 |
Yes |
4 |
380 |
GB300 |
288 |
BF16 |
Yes |
4 |
760 |
GB300 |
288 |
NVFP4 |
No |
2,4 |
240 |
GB300 |
288 |
FP8 |
No |
2,4 |
380 |
GB300 |
288 |
BF16 |
No |
4 |
760 |
B200 |
192 |
NVFP4 |
Yes |
4,8 |
240 |
B200 |
192 |
FP8 |
Yes |
8 |
380 |
B200 |
192 |
NVFP4 |
No |
2,4,8 |
240 |
B200 |
192 |
FP8 |
No |
4,8 |
380 |
B200 |
192 |
BF16 |
No |
8 |
760 |
GB200 |
192 |
NVFP4 |
Yes |
4 |
240 |
GB200 |
192 |
NVFP4 |
No |
2,4 |
240 |
GB200 |
192 |
FP8 |
No |
4 |
380 |
H200 SXM |
141 |
FP8 |
Yes |
8 |
380 |
H200 SXM |
141 |
FP8 |
No |
4,8 |
380 |
H200 SXM |
141 |
BF16 |
No |
8 |
760 |
H200 NVL |
141 |
FP8 |
Yes |
8 |
380 |
H200 NVL |
141 |
FP8 |
No |
4,8 |
380 |
H200 NVL |
141 |
BF16 |
No |
8 |
760 |
H100 SXM |
80 |
FP8 |
No |
8 |
380 |
H100 NVL |
94 |
FP8 |
No |
8 |
380 |
RTX PRO 6000 BSE |
96 |
NVFP4 |
Yes |
8 |
240 |
RTX PRO 6000 BSE |
96 |
FP8 |
No |
8 |
380 |
RTX PRO 6000 BSE |
96 |
NVFP4 |
No |
4,8 |
240 |
Qwen3.5-122B-A10B#
Latest supported release tag: 2.0.4-variant
The following section lists the supported configurations for
qwen/qwen3.5-122b-a10b (NGC catalog page).
Generic Configuration#
NIM for VLMs offers competitive performance through a custom vLLM backend. Any NVIDIA GPU with sufficient memory, or multiple homogeneous NVIDIA GPUs with sufficient aggregate memory, should be able to run this model, though this is not guaranteed. It requires compute capability >= 9.0.
The GPU Memory and Disk Space values are in GB. The disk space values for LoRA-enabled profiles do not include the LoRA model weights.
GPU |
GPU Memory |
Precision |
LoRA |
# of GPUs |
Disk Space |
|---|---|---|---|---|---|
B300 |
288 |
BF16 |
Yes |
2,4,8 |
240 |
B300 |
288 |
FP8 |
Yes |
1,2,4,8 |
120 |
B300 |
288 |
NVFP4 |
Yes |
1,2 |
80 |
B300 |
288 |
BF16 |
No |
2,4,8 |
240 |
B300 |
288 |
FP8 |
No |
1,2,4,8 |
120 |
B300 |
288 |
NVFP4 |
No |
1,2 |
80 |
GB300 |
288 |
BF16 |
Yes |
2,4 |
240 |
GB300 |
288 |
FP8 |
Yes |
1,2,4 |
120 |
GB300 |
288 |
NVFP4 |
Yes |
1,2 |
80 |
GB300 |
288 |
BF16 |
No |
2,4 |
240 |
GB300 |
288 |
FP8 |
No |
1,2,4 |
120 |
GB300 |
288 |
NVFP4 |
No |
1,2 |
80 |
B200 |
192 |
BF16 |
Yes |
2,4,8 |
240 |
B200 |
192 |
FP8 |
Yes |
1,2,4,8 |
120 |
B200 |
192 |
NVFP4 |
Yes |
1,2 |
80 |
B200 |
192 |
BF16 |
No |
2,4,8 |
240 |
B200 |
192 |
FP8 |
No |
1,2,4,8 |
120 |
B200 |
192 |
NVFP4 |
No |
1,2 |
80 |
GB200 |
192 |
BF16 |
Yes |
2,4 |
240 |
GB200 |
192 |
FP8 |
Yes |
1,2,4 |
120 |
GB200 |
192 |
NVFP4 |
Yes |
1,2 |
80 |
GB200 |
192 |
BF16 |
No |
2,4 |
240 |
GB200 |
192 |
FP8 |
No |
1,2,4 |
120 |
GB200 |
192 |
NVFP4 |
No |
1,2 |
80 |
H200 SXM |
141 |
BF16 |
Yes |
4,8 |
240 |
H200 SXM |
141 |
FP8 |
Yes |
2,4,8 |
120 |
H200 SXM |
141 |
BF16 |
No |
4,8 |
240 |
H200 SXM |
141 |
FP8 |
No |
2,4,8 |
120 |
H200 NVL |
141 |
BF16 |
Yes |
4,8 |
240 |
H200 NVL |
141 |
FP8 |
Yes |
2,4,8 |
120 |
H200 NVL |
141 |
BF16 |
No |
4,8 |
240 |
H200 NVL |
141 |
FP8 |
No |
2,4,8 |
120 |
GH200 |
141 |
FP8 |
Yes |
2 |
120 |
GH200 |
141 |
FP8 |
No |
2 |
120 |
H100 SXM |
80 |
BF16 |
Yes |
8 |
240 |
H100 SXM |
80 |
FP8 |
Yes |
4,8 |
120 |
H100 SXM |
80 |
BF16 |
No |
8 |
240 |
H100 SXM |
80 |
FP8 |
No |
4,8 |
120 |
H100 NVL |
94 |
BF16 |
Yes |
4,8 |
240 |
H100 NVL |
94 |
FP8 |
Yes |
2,4,8 |
120 |
H100 NVL |
94 |
BF16 |
No |
4,8 |
240 |
H100 NVL |
94 |
FP8 |
No |
2,4,8 |
120 |
A100 SXM |
80 |
BF16 |
Yes |
8 |
240 |
A100 SXM |
80 |
BF16 |
No |
8 |
240 |
GB10 |
128 |
NVFP4 |
No |
1 |
80 |
L40S |
48 |
BF16 |
Yes |
8 |
240 |
L40S |
48 |
FP8 |
Yes |
8 |
120 |
L40S |
48 |
BF16 |
No |
8 |
240 |
L40S |
48 |
FP8 |
No |
8 |
120 |
RTX PRO 4500 BSE |
48 |
FP8 |
Yes |
8 |
120 |
RTX PRO 4500 BSE |
48 |
FP8 |
No |
8 |
120 |
RTX PRO 6000 BSE |
96 |
BF16 |
Yes |
4,8 |
240 |
RTX PRO 6000 BSE |
96 |
FP8 |
Yes |
2,4,8 |
120 |
RTX PRO 6000 BSE |
96 |
NVFP4 |
Yes |
2 |
80 |
RTX PRO 6000 BSE |
96 |
BF16 |
No |
4,8 |
240 |
RTX PRO 6000 BSE |
96 |
FP8 |
No |
2,4,8 |
120 |
RTX PRO 6000 BSE |
96 |
NVFP4 |
No |
2 |
80 |
Kimi-K2.6#
Latest supported release tag: 2.0.4-variant
The following section lists the supported configurations for
moonshotai/kimi-k2.6
(NGC catalog page).
Generic Configuration#
Kimi-K2.6 is a ~1T-parameter MoE VLM. It ships as INT4 only, with two tensor-parallel profiles selected automatically based on per-GPU memory.
The GPU Memory column is per-GPU HBM in GB; the Disk Space column is the NGC artifact size needed in the NIM cache (one-time download on first launch), in GB.
GPU |
GPU Memory |
Precision |
# of GPUs |
Disk Space |
|---|---|---|---|---|
H200 |
141 |
INT4 |
8 |
600 |
H200 NVL |
141 |
INT4 |
8 |
600 |
B200 |
192 |
INT4 |
8 |
600 |
B300 SXM (288 GB) |
288 |
INT4 |
4 or 8 |
600 |
GB300 NVL |
288 |
INT4 |
4 |
600 |
Nemotron-3-Nano-Omni-30B-A3B-Reasoning#
Latest supported release tag: 2.0.4-variant
The following section lists the supported configurations for
nvidia/nemotron-3-nano-omni-30b-a3b-reasoning
(NGC catalog page).
Generic Configuration#
NIM for VLMs offers competitive performance through a custom vLLM backend. Any NVIDIA GPU with sufficient memory should be able to run this model, though this is not guaranteed.
The GPU Memory and Disk Space values are in GB.
GPU |
GPU Memory |
Precision |
# of GPUs |
Disk Space |
|---|---|---|---|---|
A100 SXM4 (40 GB) |
40 |
BF16 |
2 |
80 |
A100 SXM4 (80 GB) |
80 |
BF16 |
1 or 2 |
80 |
L40S |
48 |
BF16 |
2 |
80 |
L40S |
48 |
FP8 |
1 |
50 |
H100 SXM (80 GB) |
80 |
BF16 |
1 or 2 |
80 |
H100 SXM (80 GB) |
80 |
FP8 |
1 |
50 |
H100 NVL |
94 |
BF16 |
1 or 2 |
80 |
H100 NVL |
94 |
FP8 |
1 |
50 |
H200 |
141 |
BF16 |
1 or 2 |
80 |
H200 |
141 |
FP8 |
1 |
50 |
H200 NVL |
141 |
BF16 |
1 or 2 |
80 |
H200 NVL |
141 |
FP8 |
1 |
50 |
GH200 (144 GB) |
144 |
BF16 |
1 or 2 |
80 |
GH200 (144 GB) |
144 |
FP8 |
1 |
50 |
GH200 (480 GB) |
480 |
BF16 |
1 |
80 |
GH200 (480 GB) |
480 |
FP8 |
1 |
50 |
B200 |
192 |
BF16 |
1 or 2 |
80 |
B200 |
192 |
FP8 |
1 |
50 |
B200 |
192 |
NVFP4 |
1 |
40 |
GB200 NVL |
192 |
BF16 |
1 or 2 |
80 |
GB200 NVL |
192 |
FP8 |
1 |
50 |
GB200 NVL |
192 |
NVFP4 |
1 |
40 |
RTX PRO 4500 Blackwell SE |
32 |
NVFP4 |
1 |
40 |
RTX PRO 6000 Blackwell SE |
96 |
BF16 |
1 or 2 |
80 |
RTX PRO 6000 Blackwell SE |
96 |
FP8 |
1 |
50 |
RTX PRO 6000 Blackwell SE |
96 |
NVFP4 |
1 |
40 |
B300 SXM6 (288 GB) |
288 |
BF16 |
1 or 2 |
80 |
B300 SXM6 (288 GB) |
288 |
FP8 |
1 |
50 |
B300 SXM6 (288 GB) |
288 |
NVFP4 |
1 |
40 |
GB300 NVL |
288 |
BF16 |
1 or 2 |
80 |
GB300 NVL |
288 |
FP8 |
1 |
50 |
GB300 NVL |
288 |
NVFP4 |
1 |
40 |
Supported codecs and video formats#
This model supports the following codecs and video formats for an input video:
Supported codecs: H264, H265, VP8, VP9
Supported video formats: MP4, FLV, 3GP