Models#
About Model Profiles#
The models for NVIDIA NIM microservices use model engines that are tuned for specific NVIDIA GPU models, number of GPUs, precision, and so on. NVIDIA produces model engines for several popular combinations and these are referred to as model profiles. Each model profile is identified by a unique 64-character string of hexadecimal digits that is referred to as a profile ID.
The available model profiles are stored in a file in the NIM container file system.
The file is referred to as the model manifest file and the default path is /opt/nim/etc/default/model_manifest.yaml in the container.
FLUX.1-dev Model Profiles#
FLUX.1-dev is a collection of generative image AI models creating high quality, realistic images. FLUX.1-dev generates images from simple text prompts, while FLUX.1-Depth-dev and FLUX.1-Canny-dev enable greater control by combining the text prompt with an image input to guide the output image structure.
GPU |
Backend |
NIM Version Supported |
Resolution |
Variant |
Precision |
Model Profile ID |
|---|---|---|---|---|---|---|
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base |
FP4 |
1b2d236d5fa4e0425e80ff17c9480ed73f2a66a5190a102299b3c9b8936670ff |
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
canny |
FP4 |
8b42564dd5dc5dc021b47027fc25e8de3c3f20541b06643b80143facd338480b |
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
depth |
FP4 |
66188a8ebcad93374ef35c7fb89df3db16ea9176aee3515ad1a4d333d9fc8676 |
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+canny+depth |
FP4 |
b44b6dbfc4d414f5b2d11c401606380d616939bf4f9470de78b9e25de6f143e3 |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base |
FP4 |
365d6883d978bb2f2c00f5af2678115e0d92c2d09f1fe4f8bcdd813b8d731a5f |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
canny |
FP4 |
36c44753a9a188e8a36e717c4cd2d08c7c8cc4281f59c750cfda49bd9e72a0bf |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
depth |
FP4 |
387b0d749f1f6c39f7dd9b57e1e6872f809c6bf0422c71cda164be32c0fb7d79 |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+canny+depth |
FP4 |
b454d497c90956b1bf546720c1df00c1888865050d72290191f36ada319ecc6c |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base |
FP8 |
9b8c05dd711ea235c7390c838a54730dd762466484996275c6b362ed3c87d4f7 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
canny |
FP8 |
6a4f28dc7ce68a6f63cf4361cbe84341932d2c61acd6725e08fe222725be53b3 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
depth |
FP8 |
e8bf15bd38e3766339517218899a9a0ec63f4ca9d6d7086f99115b617dcf71f2 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
5ec1a6c7284f4e55127ffdceae12684c0a50242cfdeff940f53e359dc636b267 |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base |
FP8 |
9ed3f545f2316939af1984fb703115e9b706f8c7e9b4eb452f37f86a06df5bbc |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
canny |
FP8 |
cc19715f2bd209a45773ec4131c346b4c88b44d3e8f67145e719d63f6bf512d4 |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
depth |
FP8 |
a02d1b01eb43980224ebc91a471d415be2886849bce69374e9c2a63289d8debe |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
cf766e0c4e718cccf1e771e27d7bb8181120ea21219533f8d9d166f1df1bbedd |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base |
FP4 |
d3bf029a15e32df752d7bd3cad2032b4ed56b1ecc9dc780645a7f66d1a2b4776 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
canny |
FP4 |
3cdb7324f293f28218a80a1630aaa12353dceb2480143787b92e09892f59f081 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
depth |
FP4 |
9195175480a60bcf74164bce1d67858daaa0438271504c1b805220a22eb75fcf |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base+canny+depth |
FP4 |
cfdf4bb55821f71046ccdac47dc6d839b3d6732dc2f8ad6b2ee0db710b3b4377 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base |
FP4 |
5e225127aa1a4dff5ac630fea2be767e0740407d2bdea5b046892cec6083810a |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
canny |
FP4 |
52e35008640a749b88329f44104f02945ae840eb8a94d0cacbfc76b29fc6b95c |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
depth |
FP4 |
eba86f1d04eef9a139b677b34a5a799e3f84e37ce88ce9e01b5af92ce01bc268 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base+canny+depth |
FP4 |
c1cb88aa074a1d3fa8bd0bf26b67e770e9e3dfa7576807297265f120937aa1f8 |
DGX Spark |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base |
FP4 |
abaf49b800fe505242da7a6e4349cd96bed587da88a2f48b4f70a6130bcad63a |
DGX Spark |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
canny |
FP4 |
5984f523465fd200ae81849b5cecf25a377d403aaffd49957c70f421ae47a892 |
DGX Spark |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
depth |
FP4 |
6f18a167ec23b3f58d759f1819c897eb428231a770652ec907d58d63cdb5f142 |
DGX Spark |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base+canny+depth |
FP4 |
20fcca8ba47876d48bbee71b1d369cb4d8972ef60e5c805ab354f21cceec654e |
GH200 |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base |
FP8 |
bf784da42fee9d538aeda8d9430509354779a359119443942906477707f93473 |
GH200 |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
canny |
FP8 |
649eaff4b9ec9e26bca7d798f677d0c646b9de3bf932231728560fa443bc6116 |
GH200 |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
depth |
FP8 |
0255e12cc1dee6d4c7699026cea531eb09cfdb5448e0875eecd4663c8e186d95 |
GH200 |
TensorRT |
1.2.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
ac248ff9c6e6150d55975c7cc652364b1a7bb69b074dbfa748163144dc478467 |
H100 SXM |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base |
FP8 |
0376eb85528b177c914b3a435c6d34456f1ce16bd9287c7e9f22392d87de0441 |
H100 SXM |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
canny |
FP8 |
ea523d996ab2f281ca305f7de7f36f348f8203a8fe72e0bb7620931a50d82fb6 |
H100 SXM |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
depth |
FP8 |
2a971111162d9d9a60648fd97c3d5338501b538e017c302589b7c920fc81bde1 |
H100 SXM |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
1f9080e10c8ffc4ae59d15277171b0ee3fef9b987f9b45410920ad41f7c15cde |
L40S |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base |
FP8 |
fde1571bb1c3127b047f5e7ab37b48c893b055988473bab4fc5399874b964337 |
L40S |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
canny |
FP8 |
52035cc50f1e63c3cba7319f8e365f23e29442d11f768b6b87e11eea3de5cd38 |
L40S |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
depth |
FP8 |
1c55cd56fd15786b7729a3880defc5ef4284904f99f7bc6912c46e9620c43021 |
L40S |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
a12aa1e722ccc7f7685ea7663009cea0d02c49d38fce981f8177eaa6ad8e1341 |
GeForce RTX 5090 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base |
FP4 |
ac727b88271b5dc493e23ade2568954e0deaa1d76a2227a6670d6ed821fb9953 |
GeForce RTX 5090 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
canny |
FP4 |
1907fccdb6a42689ee3d448d6a93ca911f8674c2aa1ebc81b7d1f7db436eecc1 |
GeForce RTX 5090 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
depth |
FP4 |
e9d0786a812eda295914d5c7e4e1a9c989324912af3f73eeaa9631eda616d78f |
GeForce RTX 5090 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP4 |
9bd6fd188f53bce2eb42f11f81fadd1d11c3823c506b7a1c96b705f6c5e41b3a |
GeForce RTX 5080 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base |
FP4 |
9ad7ffac9b8260d15ab286637d444363a6899159e903e9cce3594a58be1489f9 |
GeForce RTX 5080 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
canny |
FP4 |
c8cfa63ee8cba592b3f52edefa18a5fda9e8f512ee3da8bc938a90336a0e75ea |
GeForce RTX 5080 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
depth |
FP4 |
3bdeb471bb31950b3a7a759b5dea3aeb80083fd328a2cee445463fcf79141373 |
GeForce RTX 5080 Laptop (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP4 |
b912befc35951aa88e450d4b0ff7ec9576688c44f434cec624d814f954b16c10 |
GeForce RTX 5070 TI (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base |
FP4 |
34f18766736a8842a8248cffc18881bf850d04698ab1e25fa7e9fc65fae82688 |
GeForce RTX 5070 TI (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
canny |
FP4 |
d65f03f4d849fd152ce78961bec9868652db607f7e7f8d02eeea68de9e964cfc |
GeForce RTX 5070 TI (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
depth |
FP4 |
48e32cc14e07205437fa4484893e017fe6ce7149de6ef3b935e61482cc43d3e7 |
GeForce RTX 5070 TI (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP4 |
094a53dd3d6b4a67e8ba8b215f996acb0f0114afc8b1a2503068ebd7e2dc4b67 |
GeForce RTX 4090 Laptop (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
base |
FP8 |
93bd95c143ef54b2ad47c47ac3d5742f9fff5ff80baa8228ce19a8577a75ebc8 |
GeForce RTX 4090 Laptop (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
canny |
FP8 |
75af532c3d833d82ad27fab8bb190f60fbb3a91b0cf70bea33d294a7c8ce5baf |
GeForce RTX 4090 Laptop (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
depth |
FP8 |
bdf9998149e94cdaf5221aa9baebc30f27449925da8eac6d1508cd945cdb643a |
GeForce RTX 4090 Laptop (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP8 |
61070e912036a7a9140d5e64126bd623522293ea3076c2f167d963a94c863b13 |
GeForce RTX 4080 (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
base |
FP8 |
96295541130ed46e3de0b25d1a95e7409784bb19da7b4b97cb586eac4e4ab778 |
GeForce RTX 4080 (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
canny |
FP8 |
842dee095df8d6ad5a2b8678605e677fea46882ff1eba1ecde76a186e8b0d1c5 |
GeForce RTX 4080 (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
depth |
FP8 |
b490222762872588294023feaecc384bbba054ae06256abd7b166d5e007cb764 |
GeForce RTX 4080 (Beta) |
TensorRT |
1.0.0-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP8 |
2a0ff19006f215b4dbb2266240c12d91ac6a005c402124c3ca3916141096fd0a |
GeForce RTX 5090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base |
FP4 |
c6d1fad563e06a49946adfa773b9117b0485ec7cd0640386f0a5884bb350a51a |
GeForce RTX 5090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
canny |
FP4 |
a1f563c2ce47feeff632d0306083ad45e05d268cffb080a34caf5f2ed14ebbcc |
GeForce RTX 5090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
depth |
FP4 |
9b45f1c8bb44d13e6d6067799e90f472001845bd76bbe4da9669214deda62eda |
GeForce RTX 5090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP4 |
d4ffdd037cbdb279689bc6f5cd969de4cdf2e63b47edc055413b759cc25bdcff |
GeForce RTX 4090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base |
FP8 |
2cf27ae9a70fb4d765e646530d14d26f380fb4cefe3c93555faaf2d84061e475 |
GeForce RTX 4090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
canny |
FP8 |
24f330eafd299ac785cc72f70cfb8d64ec1c15e16766e55ab570e6e97ef57d8b |
GeForce RTX 4090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
depth |
FP8 |
8964aba253650b90dc4bf8cd24e4c139ebd54518a9b546cb05cc2e2f23155a39 |
GeForce RTX 4090D (Beta) |
TensorRT |
1.0.1-1.1.0 |
768-1344x768-1344 |
base+canny+depth |
FP8 |
f3036de58626350a45af7c1d24b77bed31feb35848685870bb0690d18310c178 |
If your GPU model is not listed, you can use the below generic model profiles with Pytorch backend or create the model profile for your GPU using these instructions. Pytorch checkpoints are not quantized so they consume more GPU memory.
GPU |
Backend |
Resolution |
Variant |
Precision |
Model Profile ID |
|---|---|---|---|---|---|
Generic |
PyTorch |
768-1344x768-1344 |
base |
BF16 |
f0d0d4ac2ea5b121defa3e82a1fe82f289856cf5db49aa99e670e8851d8f0305 |
Generic |
PyTorch |
768-1344x768-1344 |
canny |
BF16 |
351a04dd6ca4e445f1ae4fe0da0190133c79ed4eedd2965e5da41cbb2b48826c |
Generic |
PyTorch |
768-1344x768-1344 |
depth |
BF16 |
7280cf728c45505c1a8def558d9c18534096c0fe9a976b138818e31b33e859b7 |
Generic |
PyTorch |
768-1344x768-1344 |
base+canny+depth |
BF16 |
f02c296542632aef64d11cbb13026c2502da2c290cc5b05f507a4922eedd1dda |
FLUX.1-schnell Model Profiles#
GPU |
Backend |
NIM Version Supported |
Resolution |
Precision |
Model Profile ID |
|---|---|---|---|---|---|
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
FP4 |
1b2d236d5fa4e0425e80ff17c9480ed73f2a66a5190a102299b3c9b8936670ff |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
FP4 |
365d6883d978bb2f2c00f5af2678115e0d92c2d09f1fe4f8bcdd813b8d731a5f |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
FP8 |
9b8c05dd711ea235c7390c838a54730dd762466484996275c6b362ed3c87d4f7 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
FP4 |
d3bf029a15e32df752d7bd3cad2032b4ed56b1ecc9dc780645a7f66d1a2b4776 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
FP4 |
5e225127aa1a4dff5ac630fea2be767e0740407d2bdea5b046892cec6083810a |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
FP8 |
9ed3f545f2316939af1984fb703115e9b706f8c7e9b4eb452f37f86a06df5bbc |
DGX Spark |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
FP4 |
abaf49b800fe505242da7a6e4349cd96bed587da88a2f48b4f70a6130bcad63a |
GH200 |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
FP8 |
bf784da42fee9d538aeda8d9430509354779a359119443942906477707f93473 |
H100 SXM |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
FP8 |
0376eb85528b177c914b3a435c6d34456f1ce16bd9287c7e9f22392d87de0441 |
L40S |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
FP8 |
fde1571bb1c3127b047f5e7ab37b48c893b055988473bab4fc5399874b964337 |
GeForce RTX 5090 Laptop (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP4 |
ac727b88271b5dc493e23ade2568954e0deaa1d76a2227a6670d6ed821fb9953 |
GeForce RTX 5080 Laptop (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP4 |
9ad7ffac9b8260d15ab286637d444363a6899159e903e9cce3594a58be1489f9 |
GeForce RTX 5070 TI (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP4 |
34f18766736a8842a8248cffc18881bf850d04698ab1e25fa7e9fc65fae82688 |
GeForce RTX 4090 Laptop (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP8 |
93bd95c143ef54b2ad47c47ac3d5742f9fff5ff80baa8228ce19a8577a75ebc8 |
GeForce RTX 4080 (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP8 |
96295541130ed46e3de0b25d1a95e7409784bb19da7b4b97cb586eac4e4ab778 |
GeForce RTX 5090D (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP4 |
c6d1fad563e06a49946adfa773b9117b0485ec7cd0640386f0a5884bb350a51a |
GeForce RTX 4090D (Beta) |
TensorRT |
1.0.0 |
768-1344x768-1344 |
FP8 |
ea9115a32e460d58aa89e79baee8fa1668305d5a74558d81ebfddb41a2fb3c28 |
If your GPU model is not listed, you can use the below generic model profiles with Pytorch backend or create the model profile for your GPU using these instructions. Pytorch checkpoints are not quantized so they consume more GPU memory.
GPU |
Backend |
Resolution |
Precision |
Model Profile ID |
|---|---|---|---|---|
Generic |
PyTorch |
768-1344x768-1344 |
BF16 |
f0d0d4ac2ea5b121defa3e82a1fe82f289856cf5db49aa99e670e8851d8f0305 |
FLUX.1-Kontext-dev Model Profiles#
GPU |
Backend |
NIM Version Supported |
Resolution |
Precision |
Model Profile ID |
|---|---|---|---|---|---|
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP4 |
a623d76701f895b250ad61eef210ad558deeca589d44d83e30b37075676f79f3 |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP4 |
bae3820c3c78f368301c5bbc9101cf96846caaa10a6a5205ba7ee742bf7c3564 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP8 |
5fae2421b77cf047fa59acdb47ae27d51c87a21caaef626acc512092c26d7837 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP4 |
53644af888d735870504a9cf7c8fee102174ee8a7d5ead4d9817f782c3209384 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP4 |
cd8e01f3b8e6fe80279a058a91bdaae21863325610e3e4df37bc15acfa599ac5 |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP8 |
284335f2acc80ef87c88c5a48c29085e7cac04e6dc15ad6f864dc6e59a22b52a |
DGX Spark |
TensorRT |
1.1.0+ |
672-1568x672-1568 |
FP4 |
bf2a467d95dd394e2f194467869c7952008bd9088f36e0e6984a2bed87a7874f |
GH200 |
TensorRT |
1.1.0+ |
672-1568x672-1568 |
FP8 |
25dfd90b4ddb8943a3e0945d8ccaaf6a3eb58a84bd0db1843e9b0c597cbec284 |
H100 SXM |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP8 |
66de937b2053d47cd7a508757fc3286c6e700815d746f8248b9c3541ed13fde5 |
L40S |
TensorRT |
1.0.0+ |
672-1568x672-1568 |
FP8 |
cf4230921dcf21f6ed9d1013e920eb4246cf693940295f248041745e68ec7a80 |
GeForce RTX 5090 Laptop (Beta) |
TensorRT |
1.0.0 |
672-1568x672-1568 |
FP4 |
3d9f14cd01d73e42ad3cb60ce8df5473de8263eb45e007528969712b0848f8c5 |
GeForce RTX 5080 Laptop (Beta) |
TensorRT |
1.0.0 |
672-1568x672-1568 |
FP4 |
17eff85032b1c40910c756eeaf465067027387c25243123a8089108c6054e375 |
GeForce RTX 5070 TI (Beta) |
TensorRT |
1.0.0 |
672-1568x672-1568 |
FP4 |
53a3284a85134f8c8876ccf20679f821ef4a08f006c5a64eb4d6009260b78804 |
GeForce RTX 4090 Laptop (Beta) |
TensorRT |
1.0.0 |
672-1568x672-1568 |
FP8 |
84e34939ba4f4b1fe0156dd5bfe65ddbe2085ac05bc6e359adecda9069064ae4 |
GeForce RTX 4080 (Beta) |
TensorRT |
1.0.0 |
672-1568x672-1568 |
FP8 |
70e514f75dd04bfd0055ba6b41264729ca6ff1e1770023e343638a07cbdce475 |
If your GPU model is not listed, use one of the generic model profiles below with the PyTorch backend, or create a custom model profile for your GPU by following these instructions. Because PyTorch checkpoints are not quantized, they consume more GPU memory.
GPU |
Backend |
Resolution |
Precision |
Model Profile ID |
|---|---|---|---|---|
Generic |
PyTorch |
672-1568x672-1568 |
BF16 |
6ca915ecc7893f828bf55d1882f7b3e85469edffac70bee357ea23269a870a40 |
FLUX.2-klein Model Profiles#
FLUX.2-klein is a 4B-parameter text-to-image and image-editing model. It is available with a single generic SGLang Diffusion profile.
GPU |
Backend |
Precision |
Model Profile ID |
|---|---|---|---|
Generic |
SGLang Diffusion |
BF16 |
ff6ec6066062e3f68458bc59cb6a3213abd2b0201cef0d53f0b351d3dd247317 |
Stable Diffusion 3.5 Large Model Profiles#
GPU |
Backend |
NIM Version Supported |
Resolution |
Variant |
Precision |
Model Profile ID |
|---|---|---|---|---|---|---|
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
0ae05e94c0c026e07097e59cfd9329c997ce7e4858859772904e4fa2bb792be2 |
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
7728a0a7cfb2775f59d9658c6fddc5455f92fd00cc2e0179f8dec47cac6eb1c8 |
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
b5443be5be4080d11358055ce0d7e0a912da9708eb3921501faf27198b385f33 |
GeForce RTX 5090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
c23a9a06457f1c3381b19f4fede161473f17e10d8fab37febb0baae8c9942dc7 |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
eec7a6c6835beeadfd4b1ca552c6e7ad2eb3799d98a868a7339d7c56535b8314 |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
cf0f2a7be63cddbc0eb6f0979fd0dc0974e5bdc4ed32c59a8a03de73c335924a |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
c0423868d4f4c377ec7defdf87f07a56b700549f268c14f606965e53359408c8 |
GeForce RTX 5080 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
8b36c66e07139086903ec8fe68a9a90d30cf4fc47c4da488004bd1c2cb5ca4e1 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
2c88f4bd290dc14be12b55fe60a693639490f055b1e8f62df52002dc4c92fba2 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
2d6a3d8a4a35464c4e877d096356f67f7d5c71182075b73c487e210b9527d75f |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
d2515fb73058abaa01e23bb12f6de0ee7e46d1b32d9dad129314b7d650292073 |
GeForce RTX 4090 (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
91de75fa09cbf9d9a640635531cb3ed6fcd8da5627090792cf204eb2eefdab6e |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
72b50bb2e933b5b09802668f8976287e168429488dcc87d8ffdcc1de7b8c39c0 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
842cab144f96016a90eba9e44d2e013aedfdf54c152651016a135d167fd142f7 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
0caf592f6b30f9c25f174e959487ad96d3ed9e8bcc030eb6997de74bf0e7cae0 |
NVIDIA RTX PRO 6000 Blackwell Workstation Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
5559e8f077ac11e096e74f7b684cbfc0495ecd54da62a3078c414d491dbf26a2 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
c753bf191a1cd14e278913bcddbe592e5669c34c9f7e0ee2b299f936e7cea546 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
1dbfde102469109a3dea93fb2f4ca0ec1b3dbc52c3c9da952cb41ea82c582ed4 |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
24a48159626621c11a0e6a73aab2413f2c7acb6d0ef6fe8757c3186faa53120f |
NVIDIA RTX PRO 6000 Blackwell Server Edition (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
c8ed0422b76b88cd0d81943b7eb5f2d2624a4fc2cac92ace3ce2899e4af26f6b |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
074c9a3c1d474819b9cf8d6f83c5b34dbb3873d3facdd56686866a49ec8e85dc |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
a91e0b9789298050fab70fb289ae287966b4ac83e9ae5dad4fd73b8fac760f79 |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
871a135255848eee4506b28c194e9da3fc76afe206267df4c5a922b7784c2335 |
NVIDIA RTX 6000 Ada Generation (Beta) |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
c746fe8dd3075d094cffeb2b5f7d599665a6b87975867277990a5075d0a4c1e1 |
DGX Spark |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base |
FP8 |
40ec881781b842c601d19c4fa4f54be5d215785afe92a8d4abb9273a6f8482b5 |
DGX Spark |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+canny |
FP8 |
c7056ed9a937a0cdbee1316297dc966ec60bee5ffad37f377d3f98a91a4e29f1 |
DGX Spark |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+depth |
FP8 |
17d9f3084d5b5346c2acd38effe2b832910eab8a3b1892ca81eda1cfe1c2b7e5 |
DGX Spark |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
51db77b26e65b4e8482b32d89e2b712bcb9785d12d611e21dbaed2d85e01364d |
GH200 |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base |
FP8 |
e44dd8b1721db9e285c3638c81faa164de2ac02c630fb025430ec2f62b2dd73e |
GH200 |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+canny |
FP8 |
743720b1d85bced704fd0f8846ba5c04d2d3eedfe21c00a7ec0f44a9fbc82d7d |
GH200 |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+depth |
FP8 |
b3b8e6fc9f2771a6d5c7c8dccb3ae24d4a4ebca8b88c2e66788d6bcc8358b2e7 |
GH200 |
TensorRT |
1.1.0+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
9daecb565758b5238cc019685418f06b1e091c770d5735fa76325714665183d8 |
A100 SXM |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base |
BF16 |
693c545b76b1d00523fc565442e113d767d8128e15674ffd970f24b13e1bfdb2 |
A100 SXM |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+canny |
BF16 |
45b4c6d2fe2be3e1fbf4d70ed6d378a4379e0c36cd9bdda53b9766cd163a16e6 |
A100 SXM |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+depth |
BF16 |
52abc4ce7424ac2edcaf1c6b4498bbd2657b444c8cc46514b2e4044a2c33657a |
A100 SXM |
TensorRT |
1.0.0+ |
768-1344x768-1344 |
base+canny+depth |
BF16 |
a5bd2e1d205c571b83f8e2ecf7ac35e29527b4c6f239f3fb8ea60787ae8c7515 |
H100 SXM |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
43b871abe0c6b8e3a0a59ae7e336e85184dc55d6f232be2e4bfe67f21f6b2c36 |
H100 SXM |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
a1ea47ade07e4d2011384440acbbb51f7a25a16f7dbac392e75f7f5befe6e90a |
H100 SXM |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
4d39df88e44a398c9d91a1dfb253c022b39a66975977c6467b99d6340506d02a |
H100 SXM |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
937486eb0578bc5c5f1ef4f096f89709d6caaf5d08a73bab8a9261e8f2916a48 |
L40S |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base |
FP8 |
ec3678c36831994bb00938c52fb44ccd401f475021165e9458e3069312696aac |
L40S |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny |
FP8 |
9878ecb0a807ef8eced83d82dbf74dc2bea91a3654c61f972257889fc6a5f559 |
L40S |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+depth |
FP8 |
dc37ffb26a382c6912e09e0c035f0db8d9218cf076d6a40fcf628f5ddd660847 |
L40S |
TensorRT |
1.0.1+ |
768-1344x768-1344 |
base+canny+depth |
FP8 |
da7c82185882535d9fb7d4a397884be15ee4cdbcb3d8ea3f8a8501b11c9d9c12 |
H100 SXM |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base |
BF16 |
f6c6df4fbaa14cb58201c9acbb344607bd0e6e5ff94ca8414d9cb0fa9885df05 |
H100 SXM |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base+canny |
BF16 |
c1f289eaa4b12bf6e3e97a9f83d1a828a3329c6d92f91ad28f903f91b2b69665 |
H100 SXM |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base+depth |
BF16 |
a8c2c467570f53442215ad7cf4c83bd281f4d319e186bfd3470ee4556cff676b |
H100 SXM |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base+canny+depth |
BF16 |
8ce9a981ce8c4310c4a04d647668af7075e1528f86195227dab31de4618afea4 |
L40S |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base |
BF16 |
7cb2b5e947e27f6f052a71c1e0aab137c2e81f18b3955cd896c3cca812b6feb0 |
L40S |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base+canny |
BF16 |
840a6ac59f9413b45c3e2bf448dbe5776121ca3044e7204f201b4ee7040efbfd |
L40S |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base+depth |
BF16 |
da4df560774cf40f867ecf1bfa851262c10df0f9516d8c9c38bfb5aeacbc4e22 |
L40S |
TensorRT |
1.0.0 |
768-1344x768-1344 |
base+canny+depth |
BF16 |
2a9061d7b3aad091acbddb74320268312d7b2fa203937c971387491d7bad107b |
If your GPU model is not listed, you can use the below generic model profiles with Pytorch backend or create the model profile for your GPU using these instructions. TensorRT provides on average 1.75x speedup for all variants.
GPU |
Backend |
Resolution |
Variant |
Precision |
Model Profile ID |
|---|---|---|---|---|---|
Generic |
PyTorch |
768-1344x768-1344 |
base |
BF16 |
8f23b2ce12d64905748147c73adbfe79fffabb0c2d9fa8dc95f4942dbb03d522 |
Generic |
PyTorch |
768-1344x768-1344 |
base+canny |
BF16 |
90b353eb2436047c674431dce2075e1b1f934b0e96d5bcc45db891e70f50c2d7 |
Generic |
PyTorch |
768-1344x768-1344 |
base+depth |
BF16 |
83803d2005e31831fd4bf4be11e49c5336ca103d046c62fcd7f15cbf991f9e80 |
Generic |
PyTorch |
768-1344x768-1344 |
base+canny+depth |
BF16 |
e6fd83b7f23171dade4ab3605d49a206b55394cd0ad8d8f474b2189b32c51521 |
TRELLIS Model Profiles#
The generic Pytorch profiles are optimized using Pytorch computational graph fusions.
GPU |
Backend |
Variant |
Precision |
Model Profile ID |
|---|---|---|---|---|
Generic |
PyTorch |
base:text |
BF16 |
b1ee4caec2d53864d7038621b6c27eea29106fd727787b5f72fad8f6b763dbf7 |
Generic |
PyTorch |
large:text |
BF16 |
8a7b5ca344c61411656710304a9b32eb80521b25d9b571a6f6571a9c4e4f4091 |
Generic |
PyTorch |
large:image |
BF16 |
c4ac2b36251be5c1cc3e6792ede219646c2c6dd83b18682d7521c318db8630a8 |
Generic |
PyTorch |
large:text+large:image |
BF16 |
7ce6d27f47b2ae9e076714b14b14f0cac86a3ecf53cbf258d47970aac76b2c74 |
Qwen-Image Model Profiles#
Qwen-Image is a family of 20B-parameter text-to-image foundation models with strong capabilities in complex text rendering and general image generation. Select a model version with -e NIM_MODEL_VERSION=<version>.
GPU |
Backend |
Version |
Precision |
Model Profile ID |
|---|---|---|---|---|
Generic |
SGLang Diffusion |
qwen-image |
BF16 |
82271eb1a54a008353c16e5c1d9d8e17e6cb944043052ffcc1d09b88a675e66d |
Generic |
SGLang Diffusion |
qwen-image-2512 |
BF16 |
1aec3118fa0621d4e2e320f95b65ca97504baa66186e8e6b575a13e5df445c7a |
Qwen-Image-Edit Model Profiles#
Qwen-Image-Edit is a family of 20B-parameter image-editing models that support semantic editing, appearance editing, and precise bilingual text editing. The family also includes Qwen-Image-Edit-NVPCB-OVSL2SL, an NVIDIA fine-tune of Qwen-Image-Edit model specialized for Omniverse-to-NVPCB solder-light style transfer for PCB inspection data augmentation. Select a model version with -e NIM_MODEL_VERSION=<version>.
GPU |
Backend |
Version |
Precision |
Model Profile ID |
|---|---|---|---|---|
Generic |
SGLang Diffusion |
qwen-image-edit |
BF16 |
60424c8f03327c38b64637b8d0684762585d1542569157e0a942fd6a8527dbaf |
Generic |
SGLang Diffusion |
qwen-image-edit-2509 |
BF16 |
7d5c6d0b9477bc7872b598373bd50c1084ec11bd157685d454b210ccb262ed52 |
Generic |
SGLang Diffusion |
qwen-image-edit-2511 |
BF16 |
3c8d97b4ada49e43dbce4e4965ea6bcea50b150fa7840799d47e55af52e596a9 |
Generic |
SGLang Diffusion |
qwen-image-edit-nvpcb-ovsl2sl |
BF16 |
9f4bc9a705c89987508dc8178505a0e97ad914cecec40c24e4cff10934219d33 |
Wan2.2 Model Profiles#
Wan2.2 is a 14B-parameter video-generation model family. It ships with two variants and three precision options, which you select at container startup by setting the NIM_MODEL_VARIANT and NIM_MODEL_PRECISION environment variables. Set NIM_MODEL_VARIANT to t2v for text-to-video or i2v for image-to-video. Set NIM_MODEL_PRECISION to bf16, fp8, or nvfp4.
GPU |
Backend |
Variant |
Precision |
Model Profile ID |
|---|---|---|---|---|
Generic |
TensorRT-LLM |
t2v |
BF16 |
83310ae8f14819f6e09119801b1702fbaa0cc5ffd7bfc6ca1dde9cdbb81bca2f |
Generic |
TensorRT-LLM |
t2v |
FP8 |
2d3d8c2dd5183c13be1b0075baca114544a574c89f29f98b642f242dab3ab528 |
Generic |
TensorRT-LLM |
t2v |
NVFP4 |
f23c954c2ba48263b4daea9bad0bdde11ebf83b06720fe57bb49ca5fe08c49bc |
Generic |
TensorRT-LLM |
i2v |
BF16 |
e38919caed06f51e1039e841a6c116b7388244b05ad62576100b38c935fb6924 |
Generic |
TensorRT-LLM |
i2v |
FP8 |
637a9fc36fb34d1c1ea8e75f3a8620779a1c4f036dbf1489285a850894017129 |
Generic |
TensorRT-LLM |
i2v |
NVFP4 |
6804489a9416df56fbc1e64b16ca12c356b225349134a5e4282446da4fc024f2 |
The container validates the requested precision against the host GPU’s compute capability (CC) before downloading any model weights:
bf16requires CC ≥ 8.0 (Ampere or newer)fp8requires CC ≥ 8.9 (Ada / Hopper or newer)nvfp4requires CC ≥ 10.0 (Blackwell or newer)
See the Precision Selection section of the support matrix for more details.
Wan2.2-Animate-2-14B Model Profiles#
Wan2.2-Animate-2-14B is an end-to-end character animation framework that processes driving videos using a redesigned Diffusion Transformer. By eliminating intermediate motion extractors, the framework generates high-fidelity motion while preserving character identity.
GPU |
Backend |
Precision |
Model Profile ID |
|---|---|---|---|
Generic |
SGLang Diffusion |
BF16 |
211239a1899c3d2121cff7b99afec8809db45905742ab641bbec3b7a3379047d |
Output Resolution#
The width and height request parameters are guidance values rather than exact output dimensions: their product (width x height) defines a target pixel budget for the generated video. The actual output resolution is computed from the reference image:
The output keeps the aspect ratio — and therefore the orientation — of the reference image, not of the requested
widthxheight.Both sides are snapped to multiples of 16 pixels.
The total pixel count stays within the requested
widthxheightbudget (at most 1,062,400 pixels; larger requests are rejected with a 422 status code).
As a result, the dimensions of the returned MP4 generally differ from the requested width and height. For example, a request with width=1280 and height=720 (a budget of 921,600 pixels) and a 1080x1920 portrait reference image produces a 720x1280 portrait video: the portrait orientation comes from the reference image, and 720x1280 is the largest multiple-of-16 size within the pixel budget.
To control the output size and orientation, crop or resize the reference image to the desired aspect ratio before submitting it, and set width and height so their product matches the pixel budget you want to spend.
Output Audio#
The generated MP4 is video-only: the NIM does not copy the driving video’s audio track into the output. To add sound, mux the driving video’s audio onto the generated clip for example by using FFmpeg on your machine. Refer to the Adding audio to the output instructions in Getting Started for details.