Support Matrix#

Hardware#

Unless specified otherwise, NVIDIA NIM for vision language models (VLMs) should, but are not guaranteed to, run on any NVIDIA GPU, provided the GPU has sufficient memory. They can also run on multiple homogeneous NVIDIA GPUs with sufficient aggregate memory and a CUDA compute capability of >= 7.0 (8.0 for bfloat16) unless otherwise specified. For more information, refer to Supported Models.

NVIDIA NIM for VLMs does not support NVIDIA Virtual GPU (vGPU) environments.

For information on the supported operating systems, drivers, and software, refer to the About Get Started page.

Supported Models#

Nemotron Parse v2.0#

Latest supported release tag: 2.0.8-variant

The following section lists the supported configurations for nvidia/nemotron-parse-v2.0 (NGC catalog page).

Generic Configuration#

NIM for VLMs serves this model through a custom vLLM backend. Any NVIDIA GPU with sufficient memory should be able to run this model, though this is not guaranteed.

The GPU Memory and Disk Space values are in GB.

GPU

GPU Memory

Precision

# of GPUs

Disk Space

H100-80GB-HBM3

80

bf16

1

3.39

A100-SXM4-80GB

80

bf16

1

3.39

A100-SXM4-40GB

40

bf16

1

3.39

B200

192

bf16

1

3.39

GB200

192

bf16

1

3.39

L40S

48

bf16

1

3.39

A10G

24

bf16

1

3.39

RTX PRO 6000

96

bf16

1

3.39

Inkling#

Latest supported release tag: 2.0.8-variant

The following section lists the supported configurations for thinkingmachines/inkling (NGC catalog page).

Generic Configuration#

NIM for VLMs serves this model through an SGLang backend.

The GPU Memory and Disk Space values are in GB.

GPU

GPU Memory

Precision

# of GPUs

Disk Space

B200

192

NVFP4

8

670