Support Matrix#
Models#
Model Name |
Model ID |
Publisher |
Variant |
Version |
|---|---|---|---|---|
FLUX.1-dev |
black-forest-labs/flux.1-dev |
Black Forest Labs |
base, canny, depth |
|
FLUX.1-Kontext-dev |
black-forest-labs/flux.1-kontext-dev |
Black Forest Labs |
base |
|
FLUX.1-schnell |
black-forest-labs/flux.1-schnell |
Black Forest Labs |
base |
|
FLUX.2-klein |
black-forest-labs/flux.2-klein-4b |
Black Forest Labs |
base |
|
Stable Diffusion 3.5 Large |
stabilityai/stable-diffusion-3.5-large |
Stability AI |
base, canny, depth |
|
TRELLIS |
microsoft/trellis |
Microsoft |
base:text, large:text, large:image |
|
Qwen-Image |
qwen/qwen-image |
Qwen |
qwen-image, qwen-image-2512 |
|
Qwen-Image-Edit |
qwen/qwen-image-edit |
Qwen |
qwen-image-edit, qwen-image-edit-2509, qwen-image-edit-2511 |
|
WAN2.2 |
wan-ai/wan2.2 |
Wan-AI |
t2v-14b, i2v-14b |
Supported Hardware#
Black Forest Labs / FLUX.1-dev#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
16GB |
40GB |
Linux/WSL2 |
x86_64 |
Recommended |
32GB |
64GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the FLUX.1-dev model on the following GPUs with the optimized TensorRT engine.
GPU |
NIM Versions with Pre-build Engines |
GPU Memory (GB) |
Precision |
|---|---|---|---|
GeForce RTX 5090 |
1.0.0+ |
32 |
FP4, FP8 |
GeForce RTX 5080 |
1.0.0+ |
16 |
FP4, FP8 |
GeForce RTX 4090 |
1.0.0+ |
24 |
FP8 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
1.2.0+ |
96 |
FP4, FP8 |
NVIDIA RTX PRO 6000 Blackwell Server Edition |
1.2.0+ |
96 |
FP4, FP8 |
NVIDIA RTX 6000 Ada Generation |
1.0.0+ |
48 |
FP8 |
DGX Spark |
1.2.0+ |
128 |
FP4, FP8 |
GH200 |
1.2.0+ |
96 |
FP8 |
H100 SXM |
1.1.0+ |
80 |
FP8 |
L40S |
1.1.0+ |
48 |
FP8 |
GeForce RTX 5090 Laptop |
1.0.1-1.1.0 |
24 |
FP4, FP8 |
GeForce RTX 5080 Laptop |
1.0.1-1.1.0 |
16 |
FP4, FP8 |
GeForce RTX 5070 TI |
1.0.1-1.1.0 |
16 |
FP4, FP8 |
GeForce RTX 4090 Laptop |
1.0.1-1.1.0 |
16 |
FP8 |
GeForce RTX 4080 Super |
1.0.1-1.1.0 |
16 |
FP8 |
GeForce RTX 4080 |
1.0.1-1.1.0 |
16 |
FP8 |
GeForce RTX 5090D |
1.0.1-1.1.0 |
32 |
FP4, FP8 |
GeForce RTX 4090D |
1.0.1-1.1.0 |
24 |
FP8 |
Black Forest Labs / FLUX.1-Kontext-dev#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
16GB |
40GB |
Linux/WSL2 |
x86_64 |
Recommended |
32GB |
64GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the FLUX.1-Kontext-dev model on the following GPUs with the optimized TensorRT engine.
GPU |
NIM Versions with Pre-build Engines |
GPU Memory (GB) |
Precision |
|---|---|---|---|
GeForce RTX 5090 |
1.0.0+ |
32 |
FP4, FP8 |
GeForce RTX 5080 |
1.0.0+ |
16 |
FP4, FP8 |
GeForce RTX 4090 |
1.0.0+ |
24 |
FP8 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
1.0.0+ |
96 |
FP4, FP8 |
NVIDIA RTX PRO 6000 Blackwell Server Edition |
1.0.0+ |
96 |
FP4, FP8 |
NVIDIA RTX 6000 Ada Generation |
1.0.0+ |
48 |
FP8 |
DGX Spark |
1.1.0+ |
128 |
FP4, FP8 |
GH200 |
1.1.0+ |
96 |
FP8 |
H100 SXM |
1.0.0+ |
80 |
FP8 |
L40S |
1.0.0+ |
48 |
FP8 |
GeForce RTX 5090 Laptop |
1.0.0 |
24 |
FP4, FP8 |
GeForce RTX 5080 Laptop |
1.0.0 |
16 |
FP4, FP8 |
GeForce RTX 5070 TI |
1.0.0 |
16 |
FP4, FP8 |
GeForce RTX 4090 Laptop |
1.0.0 |
16 |
FP8 |
GeForce RTX 4080 Super |
1.0.0 |
16 |
FP8 |
GeForce RTX 4080 |
1.0.0 |
16 |
FP8 |
Black Forest Labs / FLUX.1-schnell#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
16GB |
40GB |
Linux/WSL2 |
x86_64 |
Recommended |
32GB |
40GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the FLUX.1-schnell model on the following GPUs with the optimized TensorRT engine.
GPU |
NIM Versions with Pre-build Engines |
GPU Memory (GB) |
Precision |
|---|---|---|---|
GeForce RTX 5090 |
1.0.0+ |
32 |
FP4, FP8 |
GeForce RTX 5080 |
1.0.0+ |
16 |
FP4, FP8 |
GeForce RTX 4090 |
1.0.0+ |
24 |
FP8 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
1.1.0+ |
96 |
FP4, FP8 |
NVIDIA RTX PRO 6000 Blackwell Server Edition |
1.1.0+ |
96 |
FP4, FP8 |
NVIDIA RTX 6000 Ada Generation |
1.0.0+ |
48 |
FP8 |
DGX Spark |
1.1.0+ |
128 |
FP4, FP8 |
GH200 |
1.1.0+ |
96 |
FP8 |
H100 SXM |
1.0.0+ |
80 |
FP8 |
L40S |
1.0.0+ |
48 |
FP8 |
GeForce RTX 5090 Laptop |
1.0.0 |
24 |
FP4, FP8 |
GeForce RTX 5080 Laptop |
1.0.0 |
16 |
FP4, FP8 |
GeForce RTX 5070 TI |
1.0.0 |
16 |
FP4, FP8 |
GeForce RTX 4090 Laptop |
1.0.0 |
16 |
FP8 |
GeForce RTX 4080 Super |
1.0.0 |
16 |
FP8 |
GeForce RTX 4080 |
1.0.0 |
16 |
FP8 |
GeForce RTX 5090D |
1.0.0 |
32 |
FP4, FP8 |
GeForce RTX 4090D |
1.0.0 |
24 |
FP8 |
Black Forest Labs / FLUX.2-klein#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
48GB |
48GB |
Linux/WSL2 |
x86_64 |
Recommended |
48GB |
64GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the FLUX.2-klein model on any Ampere Architecture GPUs or newer with at least 48GB GPU Memory.
Stability AI / Stable Diffusion 3.5 Large#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
16GB |
48GB |
Linux/WSL2 |
x86_64 |
Recommended |
32GB |
64GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the Stable Diffusion 3.5 Large model on the following GPUs with the optimized TensorRT engine.
GPU |
NIM Versions with Pre-build Engines |
GPU Memory (GB) |
Precision |
|---|---|---|---|
GeForce RTX 5090 |
1.0.1+ |
32 |
FP8 |
GeForce RTX 5080 |
1.0.1+ |
16 |
FP8 |
GeForce RTX 4090 |
1.0.1+ |
24 |
FP8 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition |
1.0.1+ |
96 |
FP8 |
NVIDIA RTX PRO 6000 Blackwell Server Edition |
1.0.1+ |
96 |
FP8 |
NVIDIA RTX 6000 Ada Generation |
1.0.1+ |
48 |
FP8 |
DGX Spark |
1.1.0+ |
128 |
FP8 |
GH200 |
1.1.0+ |
144 |
FP8 |
A100 SXM |
1.0.0+ |
80 |
BF16 |
H100 SXM |
1.0.0+ |
80 |
FP8, BF16 |
L40S |
1.0.0+ |
48 |
FP8, BF16 |
Wan-AI / WAN2.2#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Recommended |
4 x 80GB |
128GB |
Linux/WSL2 |
x86_64, aarch64 |
Minimal |
80GB |
64GB |
Linux/WSL2 |
x86_64, aarch64 |
NVIDIA Virtual GPU (vGPU) technology is supported when the vGPU configuration meets the minimum GPU memory requirements.
The Wan2.2 NIM is published as a multi-architecture image and runs natively on both x86_64 (amd64) and aarch64 (arm64) hosts, including NVIDIA Grace-based systems such as GB200 and GB300.
NVIDIA supports the Wan2.2 model on GPUs based on the Hopper architecture or later with at least 80 GB of GPU memory. NVIDIA recommends a 4-GPU deployment (for example, 4× GB200) for the best performance. When more than one GPU is detected, the NIM automatically selects a multi-GPU parallelism configuration appropriate for the GPU count, delivering lower latency than single-GPU deployments.
Precision Selection#
Wan2.2 ships three precisions per variant. Set NIM_MODEL_PRECISION at container startup to choose one. The container validates the requested precision against the host GPU’s compute capability (CC) and available VRAM, and raises an error if the GPU cannot run it. If NIM_MODEL_PRECISION is not set, the container auto-selects the highest-fidelity precision the host GPU can run, in the order bf16 → fp8 → nvfp4.
Backend |
NIM_MODEL_PRECISION |
Required GPU Compute Capability |
|---|---|---|
BF16 TensorRT-LLM Visual Generation |
|
CC ≥ 8.0 (Ampere or newer) |
FP8 TensorRT-LLM Visual Generation |
|
CC ≥ 8.9 (Ada Lovelace / Hopper or newer) |
NVFP4 TensorRT-LLM Visual Generation |
|
CC ≥ 10.0 (Blackwell or newer) |
In addition to the compute-capability check, the container runs a VRAM precheck: if the host GPU does not have enough memory to run the requested precision at the detected GPU-count layout, startup fails with an error that suggests choosing a lower precision, using more GPUs, or deploying on a larger GPU. To override this VRAM precheck—for example, to probe the actual memory headroom of a configuration the gate rejects—set NIM_DISABLE_VRAM_PRECHECK=true when starting the container. The VRAM check is then downgraded to a warning; the compute-capability check above is still enforced, and the container may run out of memory during initialization if the precision genuinely does not fit.
Startup Time#
GPU Count |
Startup Time Range |
|---|---|
1× |
150 – 1320 seconds |
4× |
120 – 270 seconds |
These ranges are measured for warm startup and do not include the one-time model weights download that occurs on first server start.
Within each range, the actual startup time varies depending on:
The number and kind of GPUs.
The selected model variant (
t2vori2v).The selected model precision (
bf16,fp8, ornvfp4).The host’s disk I/O speed.
Microsoft / TRELLIS#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
12GB |
32GB |
Linux/WSL2 |
x86_64 |
Recommended |
24GB |
32GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the TRELLIS model on any Ampere Architecture GPUs or newer with at least 12GB GPU Memory.
Qwen / Qwen-Image#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
80GB |
64GB |
Linux/WSL2 |
x86_64 |
Recommended |
80GB |
128GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the Qwen-Image model on any Ampere Architecture GPUs or newer with at least 80GB GPU Memory.
Qwen / Qwen-Image-Edit#
System Requirements |
GPU Memory |
RAM |
OS |
CPU |
|---|---|---|---|---|
Minimal |
80GB |
64GB |
Linux/WSL2 |
x86_64 |
Recommended |
80GB |
128GB |
Linux/WSL2 |
x86_64 |
NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.
NVIDIA supports the Qwen-Image-Edit model on any Ampere Architecture GPUs or newer with at least 80GB GPU Memory.
Software#
NVIDIA Driver#
NVIDIA NIM for Visual Generative AI is built on top of Triton Inference Server which requires NVIDIA Driver release 570 or later.
Refer to the Release Notes for the detailed list of supported drivers.
NVIDIA Container Toolkit#
Your Docker environment must support NVIDIA GPUs. Refer to Installing the NVIDIA Container Toolkit for more information.
WSL2 Software Requirements#
A Windows 11 operating system (Build 23H2 and later) is supported via Windows Subsystem for Linux:
Minimum supported driver version is 570
Minimum linux distribution supported is Ubuntu 24.04
It is recommended to use Podman container management tools