Support Matrix#

Hardware#

Unless specified otherwise, NVIDIA NIM for vision language models (VLMs) should, but are not guaranteed to, run on any NVIDIA GPU, provided the GPU has sufficient memory. They can also run on multiple homogeneous NVIDIA GPUs with sufficient aggregate memory and a CUDA compute capability of >= 7.0 (8.0 for bfloat16) unless otherwise specified. For more information, refer to Supported Models.

NVIDIA NIM for VLMs does not support NVIDIA Virtual GPU (vGPU) environments.

For information on the supported operating systems, drivers, and software, refer to the About Get Started page.

Supported Models#

Gemma 4 31B IT#

Latest supported release tag: 2.0.10-variant

The following section lists the supported configurations for google/gemma-4-31b-it (NGC catalog page).

Generic Configuration#

The GPU Memory column is per-GPU memory in GB; the Disk Space column is the NGC artifact size needed in the NIM cache (one-time download on first launch), in GB.

GPU

GPU Memory

Precision

# of GPUs

Disk Space

B200

192

NVFP4

1

32.7

B300-SXM6

288

NVFP4

1

32.7

RTX-PRO-6000

96

NVFP4

2

32.7

H200

141

BF16

1

62.6

H100-80GB-HBM3

80

BF16

4

62.6

L40S

48

int4(W4A16)

2

23.3

Inkling#

Latest supported release tag: 2.0.10-variant

The following section lists the supported configurations for thinkingmachines/inkling (NGC catalog page).

Generic Configuration#

The GPU Memory column is per-GPU HBM in GB. The Disk Space column is the NGC artifact size needed in the NIM cache (one-time download on first launch), in GB.

GPU

GPU Memory

Precision

# of GPUs

# Nodes

Disk Space

H200

141

NVFP4

8

1

592

B200

192

NVFP4

8

1

592

B200

192

BF16

16

2

1,905

B300-SXM6-AC

288

NVFP4

8

1

592

All configurations use a tensor-parallel size of 8. The model has 8 key-value heads, so tensor parallelism cannot exceed 8. Configurations that span more than one node use pipeline parallelism in addition.