TensorRT Capabilities#
This section provides an overview of what you can do with TensorRT. It is intended to be useful to all TensorRT users.
Inference Library map#
Use this map to jump to the right starting point. It groups the Inference Library pages by task. The Inference Library Overview lists the same pages in sidebar order.
Task |
Page |
What you will find |
|---|---|---|
Capability overview |
This page (TensorRT Capabilities) |
Build and runtime model, precision support, refitting, and links to deeper topics |
Engine compatibility |
Version and hardware compatibility, checks, and cross-platform engine deployment |
|
First API walkthrough |
End-to-end ONNX import, engine build, and inference examples in each language |
|
Sample catalog |
Sample Explorer, build and run instructions, and cross-compiling |
|
Advanced configuration |
Refitting, precision control, tensor formats, engine tools, weight streaming, and multi-device inference |
|
Quantization and accuracy |
Q/DQ networks, PTQ/QAT, determinism, and numerical debugging |
|
Dynamic shapes |
Optimization profiles, shape tensors, and data-dependent outputs |
|
Custom layers and plugins |
Plugin V3 authoring, registration, and advanced patterns |
|
Control flow |
|
|
Edge and DLA |
DLA deployment, formats, and standalone mode |
|
Capture and replay |
Record and replay engine-building API sequences for debugging |
|
Transformers and LLMs |
Fused attention, KV cache, MoE, and transformer-specific optimizations |
C++ and Python APIs#
The TensorRT API has C++ and Python bindings. Use Python for interoperability with data-processing libraries such as NumPy and SciPy. Use C++ on platforms without Python bindings, including QNX, and when the deployment’s applicable compliance requirements mandate the C++ or safety API. Inference execution time is generally comparable because both bindings invoke the same TensorRT runtime.
Note
The Python API is not available for all platforms, such as QNX. For more information, refer to the Support Matrix.
The Programming Model#
TensorRT operates in two phases. In the first phase, usually performed offline, you provide TensorRT with a model definition, and TensorRT optimizes it for a target GPU. In the second phase, you use the optimized model to run inference.
The Build Phase#
The highest-level interface for the build phase of TensorRT is the Builder (C++, Python). The builder is responsible for optimizing a model and producing an Engine (C++, Python).
To build an engine, you must:
Create a network definition.
Specify a configuration for the builder.
Call the builder to create the engine.
The NetworkDefinition interface (C++, Python) defines the model. The most common way to transfer a model to TensorRT is to export it from a framework in ONNX format and use the TensorRT ONNX parser to populate the network definition. The table below summarizes the main entry paths when ONNX export alone is not the best fit:
Entry path |
When to use |
Next step |
|---|---|---|
ONNX export + ONNX parser |
Most training frameworks (PyTorch, TensorFlow, JAX, and others) |
Export to ONNX, then follow the C++ or Python API walkthrough |
Torch-TensorRT |
PyTorch models you want to compile in-process without a separate ONNX file |
For latest quantization schemes, consider using ModelOpt to add quantization and exporting ONNX |
TriPy |
Step-by-step network construction in Python |
|
MLIR-TensorRT |
TensorFlow, JAX, and other MLIR-capable frontends |
|
Native ``add…`` layer API |
Full manual control over the layer graph |
Populate |
You can also populate NetworkDefinition by adding layers one at a time using its add... methods and the interfaces for Layer (C++, Python) and Tensor (C++, Python).
Whichever way you choose, you must also define which tensors are the inputs and outputs of the network. Tensors that are not marked as outputs are considered to be transient values that the builder can optimize away. Input and output tensors must be named so that TensorRT knows how to bind the input and output buffers to the model at runtime.
The BuilderConfig interface (C++, Python) specifies how TensorRT should optimize the model. Among the configuration options available, you can control the TensorRT ability to reduce the precision of calculations, control the tradeoff between memory and runtime execution speed, and constrain the choice of CUDA kernels. Because the builder can take minutes or more to run, you can also control how the builder searches for kernels and cached search results for use in subsequent runs.
After you have a network definition and a builder configuration, you can call the builder to create the engine. The builder eliminates dead computations, folds constants, reorders, and combines operations to run more efficiently on the GPU. It also has multiple implementations for each layer with varying data formats. Then, it computes an optimal schedule to execute the model, minimizing the combined cost of kernel executions and format transforms.
The builder creates the engine in a serialized form called a plan, which can be deserialized immediately or saved to disk for later use.
Note
By default, TensorRT engines are specific to both the TensorRT version and the GPU on which they were created. An incompatible engine is rejected during deserialization; check the return value and TensorRT logger before creating an execution context. To configure an engine for forward compatibility, refer to the Version Compatibility and Hardware Compatibility sections.
The TensorRT network definition does not deep-copy parameter arrays, such as the weights for a convolution. Do not release the memory for those arrays until the build phase is complete. When importing a network using the ONNX parser, the parser owns the weights and must not be destroyed until the build phase is complete.
The builder times algorithms to determine the fastest. Running the builder in parallel with other GPU work can perturb the timings, resulting in poor optimization.
The Runtime Phase#
The highest-level interface for the execution phase of TensorRT is the Runtime (C++, Python).
When using the runtime, you will typically carry out the following steps:
Deserialize a plan to create an engine. Plans are executable artifacts; deserialize only engines you built or received through a trusted, authenticated channel. Refer to Engine Deserialization and the Trust Boundary.
Create an execution context from the engine.
Then, repeatedly:
Populate input buffers for inference.
Call
enqueueV3()on the execution context to run inference.
The Engine interface (C++, Python) represents an optimized model. You can query an engine for information about the input and output tensors of the network: the expected dimensions, data type, data format, and so on.
The ExecutionContext interface (C++, Python), created from the engine, is the main interface for invoking inference. The execution context contains all of the states associated with a particular invocation. You can have multiple contexts associated with a single engine and run them in parallel, but each context must be used by only one thread at a time. Create one context per concurrently executing thread. Refer to Thread-Safety Deny-List.
You must set up the input and output buffers in the appropriate locations when invoking inference. Depending on the nature of the data, this can be in either CPU or GPU memory. If not obvious based on your model, you can query the engine to determine which memory space to provide the buffer.
After the buffers are set up, inference can be enqueued (enqueueV3). The required kernels are enqueued on a CUDA stream, and control is returned to the application as soon as possible. Some networks require multiple control transfers between CPU and GPU, so control can take longer to return. To wait for completion of asynchronous execution, synchronize on the stream using cudaStreamSynchronize.
Plugins#
TensorRT has a Plugin interface that allows applications to implement operations that TensorRT does not support natively. While translating the network, the ONNX parser can find plugins created and registered with the TensorRT PluginRegistry.
TensorRT ships with a library of plugins. The source code for these plugins, along with additional plugin examples, is available on GitHub: TensorRT plugin.
You can also write your own plugin library and serialize it with the engine. The library is native code that is loaded during engine deserialization, so embed only libraries you built or obtained through a trusted channel. A plan that embeds a plugin library inherits the library’s trust requirements. Refer to Security Considerations.
Refer to the Extending TensorRT With Custom Layers section for more details.
Types and Precision#
Supported Types#
TensorRT supports FP32, FP16, BF16, FP8, FP4, INT64, INT32, INT8, UINT8, INT4, and BOOL data types. Refer to the TensorRT Operator documentation for the layer I/O data type specification.
FP32, FP16, BF16: unquantized floating point types.
INT8: low-precision integer type. Interpreted as a signed integer. Conversion to/from INT8 type requires an explicit Q/DQ layer.
INT4: low-precision integer type for weight compression. Layer and GPU support varies; verify the target combination in the TensorRT Operator documentation and the Support Matrix.
INT4 is used for weight-only-quantization. Requires dequantization before computing is performed.
Conversion to and from INT4 type requires an explicit Q/DQ layer.
INT4 weights are expected to be serialized by packing two elements per byte. For additional information, refer to the Quantized Weights section.
FP8: low-precision floating-point type
8-bit floating point type with 1-bit for sign, 4-bits for exponent, 3-bits for mantissa
Conversion to/from the FP8 type requires an explicit Q/DQ layer.
FP4: narrow-precision floating-point type
4-bit floating point type with 1-bit for sign, 2-bits for exponent, 1-bit for mantissa
Conversion to/from the FP4 type requires an explicit Q/DQ layer.
FP4 weights are expected to be serialized by packing two elements per byte. For additional information, refer to the Quantized Weights section.
UINT8: unsigned integer I/O type
On the GPU path, the data type is only usable as a network I/O type.
Network-level inputs in UINT8 must be converted from UINT8 to FP32 or FP16 using a
CastLayerbefore the data is used in other operations.Network-level outputs in UINT8 must be produced by a
CastLayerexplicitly inserted into the network (will only support conversions from FP32/FP16 to UINT8).UINT8 quantization is not supported on the GPU path. On DLA, UINT8 is available as a quantization type, together with asymmetric quantization using nonzero zero-point values, when the ONNX parser flag
kENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLAis set. Refer to ONNX Parser Flags for DLA.On the GPU path, the
ConstantLayerdoes not support UINT8 as an output type.
BOOL
A boolean type is used with supported layers.
E8M0
8-bit exponent-only floating point type (unsigned, no mantissa).
E8M0 is supported only as a scale type for MXFP8 quantization.
E8M0 has a wide dynamic range of \(\left[ 2^{-127},2^{127} \right]\).
E8M0 does not contain any mantissa bits, therefore, it can only represent exponents of 2 within its range.
E8M0 represents NaN as 0xFF.
E8M0 does not have a representation for Infinity.
Strong Typing vs Weak Typing#
Weak typing was removed in 11.0. For strongly typed network behavior, builder constraints, and migration guidance, refer to Strongly Typed Networks and the NVIDIA TensorRT Migration Guide.
Tensors and Data Formats#
When defining a network, TensorRT assumes that multidimensional C-style arrays represent tensors. Each layer has a specific interpretation of its inputs. For example, a 2D convolution will assume that the last three dimensions of its input are in CHW format. There is no option to use a WHC format. Refer to the TensorRT Operator documentation for how each layer interprets inputs.
While optimizing the network, TensorRT performs transformations internally, including to HWC and other more complex formats. The builder chooses internal formats as part of tactic selection; applications cannot set them directly. Applications can constrain formats at network and plugin I/O boundaries to reduce unnecessary transformations. Use NVIDIA Nsight Systems to identify reformat copies and measure their cost.
For additional information, refer to the I/O Formats section.
Dynamic Shapes#
By default, TensorRT optimizes the model based on the input shapes (batch size, image size, and so on) at which it was defined. However, the builder can be configured to adjust the input dimensions at runtime. To enable this, specify one or more instances of OptimizationProfile (C++, Python) in the builder configuration, containing a minimum and maximum shape for each input and an optimization point within that range.
TensorRT creates an optimized engine for each profile, choosing CUDA kernels that work for all shapes within the [minimum, maximum] range and are fastest for the optimization point (typically different kernels for each profile). You can then select among profiles at runtime.
For additional information, refer to the Working with Dynamic Shapes section.
DLA#
Note
In Enterprise TensorRT, DLA was not supported in 11.0 through 11.3; Enterprise TensorRT 11.4.0 restored DLA on platforms including Windows on ARM. It is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments. Refer to Working with DLA for the full platform and package breakdown.
TensorRT can use the NVIDIA Deep Learning Accelerator (DLA), a dedicated inference processor on many NVIDIA SoCs that supports a subset of the TensorRT layers. You can execute part of the network on the DLA and the rest on GPU. For layers that can run on either device, you can select the target device in the builder configuration on a per-layer basis.
On Windows on ARM, GPU fallback is not supported and every layer must be DLA-capable. Refer to DLA Supported Layers for Windows on ARM for that platform’s supported layers.
For additional information, refer to the Working with DLA section.
Updating Weights#
When building an engine, you can specify that its weights can be updated later. This is useful if you frequently update the model’s weights without changing the structure, such as in reinforcement learning or when retraining a model while retaining the same structure. Perform weight updates using the Refitter interface (C++, Python).
For additional information, refer to the Refitting an Engine section.
Streaming Weights#
TensorRT can be configured to stream the network’s weights from host memory to
device memory during network execution instead of placing them in device
memory at engine load time. This enables models with weights larger than free
GPU memory to run, but the transfer can increase latency. Weight streaming is
an opt-in feature at both build time (BuilderFlag::kWEIGHT_STREAMING) and
runtime (ICudaEngine::setWeightStreamingBudgetV2). Benchmark the deployed
workload both without streaming and at the intended budget; the cost depends
on the model, interconnect, and budget. Compute the budget from available
device memory and getStreamableWeightsSize(). Refer to
Weight Streaming for budget selection and scratch
memory requirements.
Note
Weight streaming is only supported with strongly typed networks. For additional information, refer to the Weight Streaming section.
trtexec#
The samples/trtexec directory in the GitHub repository includes a command-line wrapper tool called trtexec. trtexec allows you to use TensorRT without developing your application. The tool has three main purposes:
Benchmark networks on random or user-provided input data.
Generate serialized engines from models.
Generate a serialized timing cache from the builder (refer to Timing Cache).
For additional information, refer to the trtexec section.
Polygraphy#
Polygraphy is a toolkit designed to assist in running and debugging deep learning models in TensorRT and other frameworks. It includes a Python API and a command-line interface (CLI) built using this API.
With Polygraphy, you can:
Run inference among multiple backends, like TensorRT and ONNX-Runtime, and compare results (API, CLI).
Convert models to formats like TensorRT engines with post-training quantization (API, CLI).
View information about various types of models (CLI).
Modify ONNX models on the command line:
For additional information, refer to the Polygraphy repository.
Multi-Device Inference#
TensorRT supports scaling inference across multiple GPUs for workloads that benefit from distributed execution. Measure throughput and latency on the deployed topology because communication overhead can outweigh parallelism. Multi-device inference provides two capabilities:
DistCollective – Explicit distributed collective operations (
AllReduce,AllGather,Broadcast,Reduce,ReduceScatter,AllToAll,Gather,Scatter) using NVIDIA NCCL Supported on the NVIDIA Ampere architecture (SM 80) and later GPUs.Multi-device attention – Attention layers that split the key-value sequence across GPUs with context parallelism. Supported on the NVIDIA Blackwell architecture (SM 100) and later GPUs.
Collective operations require every participating rank. A lost, stalled, or slow rank can block progress for the replica group; deployments must monitor all ranks and recover the group together. Refer to Multi-Device Inference for setup details.
Note
Supported configurations are x86_64 on Ubuntu 24.04 or Rocky 8 with CUDA 12.9 or 13.2, and AArch64 on Ubuntu 22.04 with CUDA 13.2. Automotive, RTX, Coverity, and DLA builds are not supported. Review the multi-device known issues in the TensorRT 11.4.0 Release Notes, including NCCL communicator limitations on Windows and after optimization-profile changes. For setup instructions and feature compatibility details, refer to Multi-Device Inference.
For more information, refer to the sampleDistCollective sample.