Support Matrix#
Hardware#
Unless specified otherwise, NVIDIA NIM for vision language models (VLMs) should, but are not guaranteed to, run on any NVIDIA GPU, provided the GPU has sufficient memory. They can also run on multiple homogeneous NVIDIA GPUs with sufficient aggregate memory and a CUDA compute capability of >= 7.0 (8.0 for bfloat16) unless otherwise specified. For more information, refer to Supported Models.
NVIDIA NIM for VLMs does not support NVIDIA Virtual GPU (vGPU) environments.
For information on the supported operating systems, drivers, and software, refer to the About Get Started page.
Supported Models#
GLM-5.3-Flash#
Latest supported release tag: 2.1.2-variant
The following section lists the supported configurations for
zai-org/GLM-5.3-Flash
(NGC catalog page).
Generic Configuration#
The GPU Memory column is per-GPU HBM in GB unless otherwise specified. The Disk Space column is the NGC artifact size needed in the NIM cache (one-time download on first launch), in GB.
The DGX Spark configuration uses two nodes with one NVIDIA GB10 GPU each and 128 GB of unified system memory per node. For deployment instructions, refer to Deploy on DGX Spark.
GPU |
GPU Memory |
Precision |
Number of GPUs |
Disk Space |
|---|---|---|---|---|
Any |
>= 62.5 GiB |
FP8 |
8 |
334 |
NVIDIA-B200 |
196 |
FP8 |
8 |
334 |
NVIDIA-H200 |
141 |
FP8 |
8 |
334 |
NVIDIA-H100 |
80 |
FP8 |
8 |
334 |
NVIDIA-B200 |
196 |
FP8 |
4 |
334 |
NVIDIA-H20-3e |
141 |
FP8 |
8 |
334 |
NVIDIA GB10 (DGX Spark) |
128 |
NVFP4 |
2 |
208 |