Support Matrix#

Models#

Model Name

Model ID

Publisher

Variant

Version

FLUX.1-dev

black-forest-labs/flux.1-dev

Black Forest Labs

base, canny, depth

FLUX.1-Kontext-dev

black-forest-labs/flux.1-kontext-dev

Black Forest Labs

base

FLUX.1-schnell

black-forest-labs/flux.1-schnell

Black Forest Labs

base

FLUX.2-klein

black-forest-labs/flux.2-klein-4b

Black Forest Labs

base

Stable Diffusion 3.5 Large

stabilityai/stable-diffusion-3.5-large

Stability AI

base, canny, depth

TRELLIS

microsoft/trellis

Microsoft

base:text, large:text, large:image

Qwen-Image

qwen/qwen-image

Qwen

qwen-image, qwen-image-2512

Qwen-Image-Edit

qwen/qwen-image-edit

Qwen

qwen-image-edit, qwen-image-edit-2509, qwen-image-edit-2511

WAN2.2

wan-ai/wan2.2

Wan-AI

t2v-14b, i2v-14b

Supported Hardware#

Black Forest Labs / FLUX.1-dev#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

16GB

40GB

Linux/WSL2

x86_64

Recommended

32GB

64GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the FLUX.1-dev model on the following GPUs with the optimized TensorRT engine.

GPU

NIM Versions with Pre-build Engines

GPU Memory (GB)

Precision

GeForce RTX 5090

1.0.0+

32

FP4, FP8

GeForce RTX 5080

1.0.0+

16

FP4, FP8

GeForce RTX 4090

1.0.0+

24

FP8

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

1.2.0+

96

FP4, FP8

NVIDIA RTX PRO 6000 Blackwell Server Edition

1.2.0+

96

FP4, FP8

NVIDIA RTX 6000 Ada Generation

1.0.0+

48

FP8

DGX Spark

1.2.0+

128

FP4, FP8

GH200

1.2.0+

96

FP8

H100 SXM

1.1.0+

80

FP8

L40S

1.1.0+

48

FP8

GeForce RTX 5090 Laptop

1.0.1-1.1.0

24

FP4, FP8

GeForce RTX 5080 Laptop

1.0.1-1.1.0

16

FP4, FP8

GeForce RTX 5070 TI

1.0.1-1.1.0

16

FP4, FP8

GeForce RTX 4090 Laptop

1.0.1-1.1.0

16

FP8

GeForce RTX 4080 Super

1.0.1-1.1.0

16

FP8

GeForce RTX 4080

1.0.1-1.1.0

16

FP8

GeForce RTX 5090D

1.0.1-1.1.0

32

FP4, FP8

GeForce RTX 4090D

1.0.1-1.1.0

24

FP8

Black Forest Labs / FLUX.1-Kontext-dev#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

16GB

40GB

Linux/WSL2

x86_64

Recommended

32GB

64GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the FLUX.1-Kontext-dev model on the following GPUs with the optimized TensorRT engine.

GPU

NIM Versions with Pre-build Engines

GPU Memory (GB)

Precision

GeForce RTX 5090

1.0.0+

32

FP4, FP8

GeForce RTX 5080

1.0.0+

16

FP4, FP8

GeForce RTX 4090

1.0.0+

24

FP8

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

1.0.0+

96

FP4, FP8

NVIDIA RTX PRO 6000 Blackwell Server Edition

1.0.0+

96

FP4, FP8

NVIDIA RTX 6000 Ada Generation

1.0.0+

48

FP8

DGX Spark

1.1.0+

128

FP4, FP8

GH200

1.1.0+

96

FP8

H100 SXM

1.0.0+

80

FP8

L40S

1.0.0+

48

FP8

GeForce RTX 5090 Laptop

1.0.0

24

FP4, FP8

GeForce RTX 5080 Laptop

1.0.0

16

FP4, FP8

GeForce RTX 5070 TI

1.0.0

16

FP4, FP8

GeForce RTX 4090 Laptop

1.0.0

16

FP8

GeForce RTX 4080 Super

1.0.0

16

FP8

GeForce RTX 4080

1.0.0

16

FP8

Black Forest Labs / FLUX.1-schnell#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

16GB

40GB

Linux/WSL2

x86_64

Recommended

32GB

40GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the FLUX.1-schnell model on the following GPUs with the optimized TensorRT engine.

GPU

NIM Versions with Pre-build Engines

GPU Memory (GB)

Precision

GeForce RTX 5090

1.0.0+

32

FP4, FP8

GeForce RTX 5080

1.0.0+

16

FP4, FP8

GeForce RTX 4090

1.0.0+

24

FP8

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

1.1.0+

96

FP4, FP8

NVIDIA RTX PRO 6000 Blackwell Server Edition

1.1.0+

96

FP4, FP8

NVIDIA RTX 6000 Ada Generation

1.0.0+

48

FP8

DGX Spark

1.1.0+

128

FP4, FP8

GH200

1.1.0+

96

FP8

H100 SXM

1.0.0+

80

FP8

L40S

1.0.0+

48

FP8

GeForce RTX 5090 Laptop

1.0.0

24

FP4, FP8

GeForce RTX 5080 Laptop

1.0.0

16

FP4, FP8

GeForce RTX 5070 TI

1.0.0

16

FP4, FP8

GeForce RTX 4090 Laptop

1.0.0

16

FP8

GeForce RTX 4080 Super

1.0.0

16

FP8

GeForce RTX 4080

1.0.0

16

FP8

GeForce RTX 5090D

1.0.0

32

FP4, FP8

GeForce RTX 4090D

1.0.0

24

FP8

Black Forest Labs / FLUX.2-klein#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

48GB

48GB

Linux/WSL2

x86_64

Recommended

48GB

64GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the FLUX.2-klein model on any Ampere Architecture GPUs or newer with at least 48GB GPU Memory.

Stability AI / Stable Diffusion 3.5 Large#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

16GB

48GB

Linux/WSL2

x86_64

Recommended

32GB

64GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the Stable Diffusion 3.5 Large model on the following GPUs with the optimized TensorRT engine.

GPU

NIM Versions with Pre-build Engines

GPU Memory (GB)

Precision

GeForce RTX 5090

1.0.1+

32

FP8

GeForce RTX 5080

1.0.1+

16

FP8

GeForce RTX 4090

1.0.1+

24

FP8

NVIDIA RTX PRO 6000 Blackwell Workstation Edition

1.0.1+

96

FP8

NVIDIA RTX PRO 6000 Blackwell Server Edition

1.0.1+

96

FP8

NVIDIA RTX 6000 Ada Generation

1.0.1+

48

FP8

DGX Spark

1.1.0+

128

FP8

GH200

1.1.0+

144

FP8

A100 SXM

1.0.0+

80

BF16

H100 SXM

1.0.0+

80

FP8, BF16

L40S

1.0.0+

48

FP8, BF16

Wan-AI / WAN2.2#

System Requirements

GPU Memory

RAM

OS

CPU

Recommended

4 x 80GB

128GB

Linux/WSL2

x86_64, aarch64

Minimal

80GB

64GB

Linux/WSL2

x86_64, aarch64

NVIDIA Virtual GPU (vGPU) technology is supported when the vGPU configuration meets the minimum GPU memory requirements.

The Wan2.2 NIM is published as a multi-architecture image and runs natively on both x86_64 (amd64) and aarch64 (arm64) hosts, including NVIDIA Grace-based systems such as GB200 and GB300.

NVIDIA supports the Wan2.2 model on GPUs based on the Hopper architecture or later with at least 80 GB of GPU memory. NVIDIA recommends a 4-GPU deployment (for example, 4× GB200) for the best performance. When more than one GPU is detected, the NIM automatically selects a multi-GPU parallelism configuration appropriate for the GPU count, delivering lower latency than single-GPU deployments.

Precision Selection#

Wan2.2 ships three precisions per variant. Set NIM_MODEL_PRECISION at container startup to choose one. The container validates the requested precision against the host GPU’s compute capability (CC) and available VRAM, and raises an error if the GPU cannot run it. If NIM_MODEL_PRECISION is not set, the container auto-selects the highest-fidelity precision the host GPU can run, in the order bf16fp8nvfp4.

Backend

NIM_MODEL_PRECISION

Required GPU Compute Capability

BF16 TensorRT-LLM Visual Generation

bf16

CC ≥ 8.0 (Ampere or newer)

FP8 TensorRT-LLM Visual Generation

fp8

CC ≥ 8.9 (Ada Lovelace / Hopper or newer)

NVFP4 TensorRT-LLM Visual Generation

nvfp4

CC ≥ 10.0 (Blackwell or newer)

In addition to the compute-capability check, the container runs a VRAM precheck: if the host GPU does not have enough memory to run the requested precision at the detected GPU-count layout, startup fails with an error that suggests choosing a lower precision, using more GPUs, or deploying on a larger GPU. To override this VRAM precheck—for example, to probe the actual memory headroom of a configuration the gate rejects—set NIM_DISABLE_VRAM_PRECHECK=true when starting the container. The VRAM check is then downgraded to a warning; the compute-capability check above is still enforced, and the container may run out of memory during initialization if the precision genuinely does not fit.

Startup Time#

GPU Count

Startup Time Range

150 – 1320 seconds

120 – 270 seconds

These ranges are measured for warm startup and do not include the one-time model weights download that occurs on first server start.

Within each range, the actual startup time varies depending on:

  • The number and kind of GPUs.

  • The selected model variant (t2v or i2v).

  • The selected model precision (bf16, fp8, or nvfp4).

  • The host’s disk I/O speed.

Microsoft / TRELLIS#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

12GB

32GB

Linux/WSL2

x86_64

Recommended

24GB

32GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the TRELLIS model on any Ampere Architecture GPUs or newer with at least 12GB GPU Memory.

Qwen / Qwen-Image#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

80GB

64GB

Linux/WSL2

x86_64

Recommended

80GB

128GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the Qwen-Image model on any Ampere Architecture GPUs or newer with at least 80GB GPU Memory.

Qwen / Qwen-Image-Edit#

System Requirements

GPU Memory

RAM

OS

CPU

Minimal

80GB

64GB

Linux/WSL2

x86_64

Recommended

80GB

128GB

Linux/WSL2

x86_64

NVIDIA Virtual GPU (vGPU) technology is supported if vGPU configuration meets at least minimal GPU memory requirements.

NVIDIA supports the Qwen-Image-Edit model on any Ampere Architecture GPUs or newer with at least 80GB GPU Memory.

Software#

NVIDIA Driver#

NVIDIA NIM for Visual Generative AI is built on top of Triton Inference Server which requires NVIDIA Driver release 570 or later.

Refer to the Release Notes for the detailed list of supported drivers.

NVIDIA Container Toolkit#

Your Docker environment must support NVIDIA GPUs. Refer to Installing the NVIDIA Container Toolkit for more information.

WSL2 Software Requirements#

A Windows 11 operating system (Build 23H2 and later) is supported via Windows Subsystem for Linux:

  1. Minimum supported driver version is 570

  2. Minimum linux distribution supported is Ubuntu 24.04

  3. It is recommended to use Podman container management tools