Support Matrix#

Hardware#

Unless specified otherwise, NVIDIA NIM for vision language models (VLMs) should, but are not guaranteed to, run on any NVIDIA GPU, provided the GPU has sufficient memory. They can also run on multiple homogeneous NVIDIA GPUs with sufficient aggregate memory and a CUDA compute capability of >= 7.0 (8.0 for bfloat16) unless otherwise specified. For more information, refer to Supported Models.

NVIDIA NIM for VLMs does not support NVIDIA Virtual GPU (vGPU) environments.

For information on the supported operating systems, drivers, and software, refer to the About Get Started page.

Supported Models#

Qwen3.8-Flash-Next#

Latest supported release tag: 2.1.1

The following section lists the supported configurations for Qwen/Qwen3.8-Flash-Next using the vLLM Model-Free NIM.

Generic Configuration#

The table lists the supported GPU configurations and maximum context length.

GPU (Memory)

GPU Count

Weight Precision

vLLM dtype

Tensor Parallelism

Maximum Context Length

NVIDIA H20-3e (141 GB)

4

FP8

BF16

4

262,144

NVIDIA B200 (180 GB)

2

FP8

BF16

2

262,144

NVIDIA H200 (141 GB)

4

FP8

BF16

4

262,144

NVIDIA H20-3e (141 GB)

4

BF16

BF16

4

262,144

NVIDIA B200 (180 GB)

4

BF16

BF16

4

262,144

NVIDIA H200 (141 GB)

4

BF16

BF16

4

262,144