TensorRT 11.4.0 Release Notes#
These Release Notes apply to x86 Linux and Windows users, ARM-based CPU cores for Server Base System Architecture (SBSA) users on Linux, and Windows on ARM users. This release includes several fixes from the previous TensorRT releases and additional changes.
Announcements#
Windows on ARM: TensorRT 11.4.0 adds Windows on ARM support for systems with supported DLA hardware. TensorRT executes only on DLA on this platform; there is no GPU execution path.
TensorRT-Model-Connect is in public preview: TensorRT-Model-Connect is a separate open-source project, not a component of the TensorRT 11.4.0 packages. It provides a large collection of pre-implemented open models, taking a supported Hugging Face or local checkpoint to end-to-end TensorRT inference in two commands with no intermediate ONNX export step. As a public preview reference implementation, its APIs, scope, and direction may change. Refer to Models and Recipes for supported checkpoints and to the Quick Start to build and run one.
CUDA green-context guidance: This release documents how TensorRT applications can use CUDA green contexts to partition streaming multiprocessors and work queues. Green contexts require CUDA 13.1 or later, do not partition memory, and do not provide MIG-style strict isolation. Refer to Green Contexts.
Key Features and Enhancements#
DLA support in TensorRT 11.4.0: Enterprise TensorRT 11.4.0 restores DLA on Windows on ARM only, where DLA support is in beta. DLA was unavailable in TensorRT 11.0 through 11.3, including TensorRT 11.3.1 for DriveOS. There is no GPU fallback on Windows on ARM, so every layer must be DLA-capable, and batch size is 1. In 11.4.0, DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments, including NVIDIA DRIVE AGX Thor. Use TensorRT 10.7 if DLA is required on those platforms. Refer to DLA Supported Layers for Windows on ARM and Working with DLA.
Important
DLA support on Windows on ARM is in beta and covers a restricted set of layers. DLA engines built with TensorRT 11.4.0 may fail to load once DLA support reaches general availability, and must be rebuilt with that later TensorRT release. Keep your source models so you can rebuild those engines.
Reduced build-time memory usage with weight placeholders: TensorRT 11.3.0 introduced support for null weights in the weight-stripping workflow. TensorRT 11.4.0 reduces host and device memory use for this workflow. The reduction is network-dependent; measure peak memory on the target build workload before changing build-host capacity. For API usage details, refer to Reducing Build-Time Memory Usage with Weight Placeholders.
Breaking ABI Changes#
TensorRT 11.0 removed weak typing APIs along with several other deprecated APIs. Refer to the NVIDIA TensorRT Migration Guide for more information.
TensorRT 11.3.0 changed the safety Consistency Checker C++ API (
nvinfer2::safe::consistency) from STL-container parameters to pointer and count parameters.createConsistencyCheckernow takeschar const* const* pluginBuildLibsandint64_t nbPluginBuildLibsinstead ofstd::vector<std::string>, and the serialized-blob size argument isint64_tinstead ofsize_t.IPluginChecker::validatenow takes a pointer and countsTensorDescriptorarguments instead ofstd::vector<TensorDescriptor>. Applications that callcreateConsistencyCheckerand custom libraries that implementIPluginCheckermust rebuild against matching TensorRT 11.3.0 headers and libraries. Mixing binaries built against the previous signatures with the new checker library can crash.
Breaking Behavioral Changes#
trtexec
The
--useCudaGraph,--noDataTransfers,--useSpinWait, and--separateProfileRunflags were enabled by default and deprecated in TensorRT 11.0. Each flag is still accepted but has no effect. For replacement flags and details, refer to the trtexec migration page.The
--stronglyTypedflag has no effect in TensorRT 11.0.0+, since strongly typed networks are now the default. The flag is still accepted for backward compatibility.
Compatibility#
For comprehensive platform compatibility, hardware requirements, and feature availability information, refer to the TensorRT Support Matrix. The support matrix provides detailed information about supported operating systems, CUDA versions, GPU architectures, precision modes, compiler requirements, and ONNX operator support for TensorRT 11.4.0.
Limitations#
DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments in TensorRT 11.4.0. Use TensorRT 10.7 if DLA is required on those platforms. Refer to Working with DLA for the platforms and packages that support DLA.
NVIDIA JetPack is not supported in TensorRT 11.4.0. Jetson deployments must remain on a TensorRT 10.x release supported by their JetPack version.
Version Compatible engine builds from ONNX models with explicit quantization/dequantization (Q/DQ) nodes may fail during engine build if Q/DQ is per-channel scaling for convolution filter.
When implementing a custom layer using
IPluginV3plugin class where the custom layer has data-dependent shape (DDS), the size tensors must be of onlyINT64type and notINT32type, as the latter would result in a compilation failure. Related samples have been updated accordingly.
Deprecated API Lifetime#
APIs deprecated in TensorRT 11.4.0 will be retained until October 2027.
APIs deprecated in TensorRT 11.3.0 will be retained until August 2027.
APIs deprecated in TensorRT 11.2.1 will be retained until July 2027.
APIs deprecated in TensorRT 11.1.0 will be retained until June 2027.
APIs deprecated in TensorRT 11.0.0 will be retained until March 2027.
APIs deprecated in TensorRT 10.16 will be retained until March 2027.
APIs deprecated in TensorRT 10.15 will be retained until January 2027.
See also
- Migrating from TensorRT 10.x to 11.x
Step-by-step migration guide for upgrading from TensorRT 10.x.
- Upgrading TensorRT
Package-level upgrade instructions (pip, Debian, RPM, tar, zip).
Deprecated and Removed Features#
Windows 10 support is considered deprecated and Windows 10 will no longer be supported by TensorRT after October 2026. Windows 10 has been End-of-Life since October 2025. GeForce driver support for Windows 10 will also be reduced at that time. For more information, refer to GeForce Support Plan for Windows 10.
The TensorRT Python bindings for Python versions 3.8 and 3.9 are deprecated. These Python versions are considered End-of-Life by Python upstream and have not been supported by the TensorRT samples for multiple releases. These bindings will be removed in a future release.
The Detectron 2 Python sample (
samples/python/detectron2) is removed in TensorRT 11.3.0. Refer to the Sample Explorer for remaining samples, including object detection.
Fixed Issues#
Fixed a calibration failure in NVIDIA TensorRT Model Optimizer ONNX quantization that reported
ValueError: Too many bins for data range. Cannot create 128 finite-sized bins. The failure occurred in the ModelOpt ONNX calibration histogram path (ONNX Runtime calibration callingnp.histogram), not during TensorRT engine build or inference. It was observed for some explicitly quantized ONNX models, including networks targeting the NVIDIA Blackwell architecture and quantized ResNet-50 transformer configurations, on SBSA platforms with CUDA 13.3. The fix is delivered in NVIDIA TensorRT Model Optimizer rather than in TensorRT.Fixed incorrect numerical output from engines built with
kREFIT, which could differ from an equivalent non-refit engine even when no refit was performed. A GEMM bias could be dropped and treated as zero, and because that bias was not exposed through the refit API the discrepancy could not be corrected by refitting. Both engines were deterministic across runs, and engines built withkREFIT_IDENTICALwere not affected.Fixed GPU compute time regressions on NVIDIA DGX Spark (compute capability 12.1) where some networks ran slower in TensorRT 11.3.0 than in TensorRT 11.2.1. Reproduced cases included 3dUNet INT8 (about 43% to 44% higher) and hardware-compatible BERT INT8 (about 23% higher). These were independent regressions: the 3dUNet INT8 slowdown was associated with the CUDA 13.4 device compiler, and the hardware-compatible BERT INT8 slowdown with a slower GELU path on compute capability 12.1.
Fixed a GPU time regression for strongly typed FP16 DeBERTa, which ran slower in TensorRT 11.3.0 than in TensorRT 11.2.1 because fused multi-head attention was lost and replaced by extra GEMM and memory-movement work. Observed increases were about 16.7% on NVIDIA RTX 4090 (about 12.25 ms to 14.30 ms) and NVIDIA L40S, and about 20% on NVIDIA DGX Spark, with activation memory growing from about 105 MB to about 197 MB.
Fixed a GPU time regression for a strongly typed FP16 dense network on NVIDIA B200, which ran about 8.4% slower in TensorRT 11.3.0 than in TensorRT 11.2.1, from about 1.68 ms to 1.82 ms, and used 120 MiB more execution-context GPU memory, from about 839 MiB to 959 MiB. The slowdown was associated with a change in GEMM fusion, where separate GEMMs were replaced by fused MatMul work plus extra concatenation and memory movement.
Fixed an engine build failure for some strongly typed ONNX networks that use float quantization, which failed during tactic cost computation with
Could not find any implementation for node {ForeignNode[...]}. The failure was observed on NVIDIA GH200 480 GB SBSA systems running Ubuntu 22.04, including quantized encoder and RoBERTa-base Q/DQ models built withtrtexec.Fixed an engine build failure for a strongly typed LSTM autoencoder ONNX network with dynamic batch and sequence length 10. Shape verification failed on a reshape whose element count did not match, and a follow-on
Could not find any implementation for node {ForeignNode[...]}error could appear after that check failed. The failure was observed on NVIDIA RTX PRO 5000 Blackwell AArch64 systems running Ubuntu 26.04.Fixed an execution context creation failure for a
GridSampleFP16 network on Windows with NVIDIA RTX 5080, whereICudaEngine::createExecutionContextreturned TensorRT error code 1 with CUDA error 700 while loading a CUDA module. The failure was observed in TensorRT 11.3.0 on a case that previously passed on the same GPU with TensorRT 11.2.
Known Issues#
Functional
On CUDA versions prior to 13.2, the compute-sanitizer
initchecktool may flagUninitialized __global__ memory readerrors when running TensorRT applications on NVIDIA Hopper GPUs. These errors are false positives and can be safely ignored. To suppress them, upgrade to CUDA 13.2 or later.Running TensorRT applications under Valgrind memcheck on CUDA 13.3 may report host-side memory leaks attributed to NVRTC components used during engine build and runtime compilation. These reports are not known to affect normal inference behavior. This will be addressed in a future CUDA or TensorRT release.
Running TensorRT inference under compute-sanitizer
memcheckmay report device memory errors during execution or engine teardown. These reports are not known to affect normal inference behavior when sanitizers are not enabled. This will be addressed in a future release.On NVIDIA B100, networks that use multi-head attention with packed query and variable-length key-value input (
IAttentionwithsetKeyValueLengths) may fail under compute-sanitizermemcheck. The sanitizer can report invalid__global__atomic accesses in the attention kernel, followed by CUDA error 719 during memory deallocation and an unrecoverable CUDA error. This is not a memory leak; sanitizer leak summaries showed 0 leaked bytes. Equivalent runs without the sanitizer have passed. A related failure on NVIDIA B100 may reject an unsupported stride order[3, 2, 1, 0]for a key operand under compute-sanitizer or Valgrind. This will be addressed in a future release.inplace_addmini-sample of thequickly_deployable_pluginsPython sample may produce incorrect outputs on Windows.TensorRT may exit if inputs with invalid values are provided to the
RoiAlignplugin (ROIAlign_TRT), especially if there is inconsistency in the indices specified in thebatch_indicesinput and the actual batch size used.On Windows, Version Compatible engines built with TensorRT 10.1 through 10.4 that use the
RoiAlignplugin (ROIAlign_TRT) cannot be deserialized with the TensorRT 11.x runtime due to a bug in how those engines were produced. Deserialization fails inside the runtime dispatch VC plugin loader when loadingnvinfer_vc_plugin.dll(WindowsERROR_MOD_NOT_FOUND/ error 126), surfaces as TensorRT error code 6 (API Usage Error), and leaves the returnedICudaEnginepointer null. The bug was fixed in TensorRT 10.5. Engines built with TensorRT 10.5 or later are expected to deserialize correctly. As a workaround, rebuild affected engines with TensorRT 10.5 or later before loading them with TensorRT 11.x.Creating an execution context for a Version Compatible engine may terminate the process with a segmentation fault. The engine deserializes successfully, and the crash occurs while the lean runtime initializes the execution context and resolves a matrix-multiplication operation. The failure has been observed with
trtexec --vcfor strongly typed FP16 ResNet-50 v1.5 and ViT networks on CUDA 13.2, and the same networks run correctly when the engine is built without--vc. As a workaround, build the engine without version compatibility. This will be addressed in a future release.Inspecting very large engines with the Engine Inspector may fail on systems with limited host memory. The inspection process can exceed available memory and be terminated by the operating system without a TensorRT error message. This has been observed with very large diffusion-model engines. As a workaround, run engine inspection on a system with sufficient host memory.
Building TensorRT OSS samples with Visual Studio / MSBuild on Windows may fail during the CMake configure step when
CUDA_HOMEandCUDA_PATHpoint to different CUDA Toolkit installations. The failure is environment-wide rather than sample-specific and can affect multiple sample targets andtrtexec. As a workaround, ensureCUDA_HOMEandCUDA_PATHrefer to the same CUDA Toolkit, or unset one of the variables so CMake discovers a single consistent toolkit. This will be addressed in a future release.Building an SDXL UNetXL ONNX network with
trtexecmay fail during optimizer shape analysis with an API Usage Error fromIConcatenationLayer. The failure has been observed on Windows with NVIDIA RTX 4090 for standard and strongly typed FP16 UNetXL ONNX models, with and without TF32. The error is raised at the/up_blocks.1/Concatnode and reports that axis 2 dimensions must be equal for concatenation on axis 1, with mismatched dimensions 2 and 1. The failure occurs whentrtexecauto-overrides the sample input shape to1x4x1x1, which produces mismatched spatial dimensions on the UNet down and up path before concatenation. As a workaround, specify the intended input shape sotrtexecdoes not auto-override it to1x4x1x1. This will be addressed in a future release.Building a FLUX.1-dev single-transformer attention ONNX network with
trtexecmay fail during optimizer shape analysis with an API Usage Error fromIAttentionInputLayer. The failure has been observed on Windows with NVIDIA RTX 4090, with and without TF32. The error is raised atfused_attention_0_AttentionInputand reports that Key and Value do not match at dimension 2. Before the check fails,trtexecmay auto-override unspecified dynamic inputs, includinghiddento1x1x3072,img_idsto1x3, andpooled_projectionsto1x768. This will be addressed in a future release.IRuntime::getEngineValiditymay reportEngineValidity::kVALIDfor engine data that was modified after serialization, rather thanEngineValidity::kINVALID. The failure has been observed on Windows with NVIDIA RTX 4090. Do not rely on this API alone to detect a tampered or truncated engine plan. This will be addressed in a future release.Changing input shapes on a network that uses an
IPluginV3plugin may fail with TensorRT error code 2 (Internal Error) from the plugin shape-change path. The failure has been observed as a regression on NVIDIA A100 with CUDA 12.6. This will be addressed in a future release.Multi-device inference may fail to initialize its NCCL communicator on Windows platforms, so
enqueueV3returns TensorRT error code 1 and distributed collective operations such asall_gatherdo not run. The failure has been observed with the distributed collective sample on Windows with eight NVIDIA A10 GPUs and CUDA 12.9. Another multi-device related problem was observed on Linux with eight NVIDIA A30 GPUs, where the NCCL communicator was not usable after setting a new optimization profile These will be addressed in a future release.Running inference with
trtexecmay fail with a CUDA unspecified launch failure on NVIDIA Blackwell architecture GPUs for networks whose engine compiles into a single fused subgraph. Observed cases include INT8 per-channel quantized BERT QAT models and strongly typed FP16 NVIDIA Maxine™ speech models on Linux with CUDA 13.4, and some runs instead reach the test time limit rather than crashing. This will be addressed in a future release.Enqueuing an engine that contains a data-dependent-shape operation while the CUDA stream is in graph-capture mode may fail with TensorRT error code 1 and CUDA error 900 (
operation not permitted when stream is capturing). TensorRT first logs a warning that the engine contains operations that are not permitted under CUDA graph capture mode. The failure can also leave the stream and the allocator in an inconsistent state, producing a follow-on deallocation error. Do not capture engines with data-dependent shapes into a CUDA graph. This will be addressed in a future release.The Engine Inspector may omit an original ONNX layer name from the per-layer metadata of a built engine. A name set on a transpose operation can be dropped when graph optimization folds that operation into a view, so the originating ONNX layer cannot be found in the inspector output. The engine builds and runs correctly, and only layer attribution is affected. The behavior has been observed with a quantized ResNet-50 Q/DQ model on NVIDIA RTX 5090 and on NVIDIA DGX Spark with CUDA 13.2 and later. This will be addressed in a future release.
A strongly typed post-training quantized INT8 plus FP16 YOLOX-Tiny network at 640 resolution may produce numerically incorrect output. The observed median error requires a tolerance of about 0.43 absolute or about 9.6% relative, against an expected tolerance of 0.05 absolute or 2% relative. The failure has been observed on Windows with NVIDIA RTX 5080 on CUDA 12.9 and NVIDIA RTX 5090 on CUDA 13.0, and the same network passed in TensorRT 11.3.0. This will be addressed in a future release.
On Windows on ARM, building a DLA network that contains a constant with no consumers may fail during engine build with an internal assertion from the DLA constant path. Such a constant serves no purpose and is unlikely to appear in a real network. As a workaround, remove any constant that no layer consumes. This will be addressed in a future release.
On Windows on ARM, DLA networks that use
UINT8tensors may fail during engine build. Reported cases include a constant that produces aUINT8output, a network-levelUINT8output that is not produced byIIdentityLayerorICastLayer, a strongly typed network whose requestedUINT8output type does not match the inferredINT8type, and aDequantizeLinearnode with aUINT8input. As a workaround, insert anICastLayerto convertUINT8tensors to a supported type. This will be addressed in a future release.On Windows on ARM, DLA does not support
BOOLconstants. A network that places aBOOLconstant on the DLA path fails during engine build with TensorRT error code 1 reporting an unsupported data type. Because Windows on ARM has no GPU fallback, the build cannot place the affected layer on the GPU instead. This restriction applies to constants rather than to theBOOLtype in general. An ElementWise comparison can run on DLA when it matches one of two patterns:Equal,Greater, orLessfollowed by aCastto FP16 or INT8; orGreaterorLessfollowed byNotand then aCastto FP16 or INT8. A comparison that does not match one of these patterns fails the build with a message that lists them. This will be addressed in a future release.On Windows on ARM, quantized DLA networks may fail during engine build on DLA quantization constraints. Reported cases span the
Add,Slice,Expand,GridSample,GlobalAveragePool, andInstanceNormalizationoperators. A failure reports that a node cannot be quantized by its input and that a dequantize node is needed before it. As general guidance, follow every quantize (Q) node with a dequantize (DQ) node so that each quantized value is consumed by a DQ node. This will be addressed in a future release.On Windows on ARM, a cast that TensorRT inserts on the DLA path may fail the build with TensorRT error code 4, reporting that the layer is not supported on DLA while GPU fallback is not enabled. The ONNX parser inserts such a cast for operators including
TopKandGather, and a cast in the postprocessing part of a network can fail the same way. A standaloneCastmay also be unplaceable on DLA under the numeric-safety constraints that DLA applies. Windows on ARM does not support GPU fallback, so every layer must be DLA-capable. As a workaround, remove the cast from the ONNX model, or keep the affected subgraph off the DLA path. This will be addressed in a future release.On Windows on ARM, a strongly typed DLA network may show an intermittent accuracy drop. The drop has been observed when a model runs as part of a larger suite and does not reproduce when the same model is run on its own, so a single re-run is not sufficient to confirm accuracy. This will be addressed in a future release.
Performance
On NVIDIA RTX PRO 6000 Blackwell (compute capability 12.0), a strongly typed FP16 ViT-Base patch-16 384 network (
vit_base_patch16_384_strongType_fp16.onnx) may show about 6% to 7% higher GPU time in TensorRT 11.3.0 than in TensorRT 11.2.1. The slowdown is associated with tactic selection for the MLPfc2MatMul layers and does not reproduce on NVIDIA Ada Lovelace architecture GPUs. This will be addressed in a future release.On NVIDIA GH200 480 GB SBSA (AArch64), FP8 causal attention may show about 11% higher GPU compute time in TensorRT 11.3.0 than in TensorRT 11.2.1. The slowdown is associated with the CUDA 13.4 device compiler: fused multi-head attention kernel source is unchanged, but the compiled device code is slower. Engine structure and tactic selection are otherwise the same. This will be addressed in a future release.
A non-zero
tilingOptimizationLevelmight introduce engine build failures for some networks on L4 GPUs.Engines built with
kREFITorkREFIT_IDENTICALhave performance regressions compared with non-refit engines where convolution layers are present within a branch or loop, and the precision is FP16/INT8.A strongly typed BF16 BERT base cased network built as a refittable engine may show higher inference time in TensorRT 11.4.0 than in the previous release. The regression has been observed on NVIDIA H100 with CUDA 13.4 at sequence length 512. This will be addressed in a future release.
Some Maxine™ AR Facial Landmark Detection models (
landmark_quality_68_dynamic_strongtype_fp16andlandmark_quality_126_dynamic_strongtype_fp16) may exhibit approximately 7–9% higher GPU inference time in TensorRT 11.2.x than in TensorRT 11.1.0 on NVIDIA RTX PRO 6000 Blackwell Max-Q. It is for functional fixing, absolute impact is typically about 0.1 ms on workloads of roughly 1.7–1.9 ms.