FAQs#

This section provides answers to the most frequently asked questions about TensorRT, organized by topic.

Why is trtexec not found after installation?

Pip wheels do not include trtexec. Use a Debian, RPM, tar, zip, or container installation for CLI workflows. For Debian and RPM packages, use the trtexec_path verification command on the Debian or RPM page to locate the executable and add its directory to PATH. Tar and zip users must complete the PATH step on their installation page.

What is an engine, and why must I build one?

A TensorRT engine, also called a plan, is the optimized executable artifact produced from a model for a specific deployment configuration. Building selects implementations and records GPU-specific executable code. Refer to Build Your First Engine for the shortest build-and-run workflow. Deserialize only engines you built or received through a trusted channel.

What should I do when my first engine build fails?

Retain the complete trtexec output. Confirm the GPU, driver, CUDA package, and platform in the Support Matrix, then match the error text against Troubleshooting. A compatibility match does not diagnose the failure by itself.

How do I create an optimized engine for several batch sizes?

While TensorRT allows an engine optimized for a given batch size to run at any smaller size, the performance for those smaller sizes cannot be as well optimized. To optimize for multiple batch sizes, create optimization profiles at the dimensions assigned to OptProfileSelector::kOPT. Refer to Optimization Profiles.

Are engines portable across TensorRT versions?

By default, no. Deserialization can fail when the TensorRT version, GPU, or configured compatibility mode does not match the engine. Refer to Version Compatibility for supported compatibility combinations and required build settings.

How do I choose the optimal workspace size?

Some TensorRT algorithms require additional workspace on the GPU. The method IBuilderConfig::setMemoryPoolLimit() controls the maximum workspace and prevents algorithms that require more from being considered. Set a limit that leaves memory for the model, build process, display or peer workloads, and operating-system overhead. An unbounded build can exhaust device or host memory and can be terminated by the operating system. At runtime, TensorRT allocates no more workspace than required, even if the configured limit is higher. Refer to Bounding TensorRT Memory for memory monitoring and bounding guidance.

How do I use TensorRT on multiple GPUs?

TensorRT supports parallelizing workloads across multiple GPUs. For more information, refer to the Inference Library Overview.

You may also use TensorRT on a single, specific GPU when multiple are available. Each ICudaEngine object is bound to a specific GPU when it is instantiated, either by the builder or on deserialization. The CUDA current device is per-thread. Every worker thread must call cudaSetDevice() before building, deserializing, or using an execution context bound to that device. Each IExecutionContext is bound to the same GPU as its engine.

How do I get the version of TensorRT from the library file?

The library exports a symbol whose name encodes the TensorRT version. On Linux, list the matching symbol with:

nm -D libnvinfer.so.* | grep 'tensorrt_version_'

Read the numeric suffix from the command’s actual output. To query the installed Python package instead, run python3 -c "import tensorrt as trt; print(trt.__version__)".

What can I do if my network produces the wrong answer?

There are several reasons why your network can be generating incorrect answers. Here are some troubleshooting approaches that can help diagnose the problem:

  • Turn on VERBOSE-level messages from the log stream and check what TensorRT is reporting.

  • Check that your input preprocessing generates exactly the input format the network requires.

  • If you built with BuilderFlag::kREFIT, compare the engine against an equivalent non-refit engine to establish whether refitting is involved at all. Refer to Refitting an Engine.

  • If you are using reduced precision, run the network in FP32. If it produces the correct result, lower precision can have an insufficient dynamic range for the network.

  • Try marking intermediate tensors in the network as outputs and verify if they match your expectations.

  • Use NVIDIA Nsight Deep Learning Designer to inspect your compiled engine.

Note

Marking tensors as outputs can inhibit optimizations and, therefore, can change the results.

You can use NVIDIA Polygraphy to assist you with debugging and diagnosis.

How do I implement batch normalization in TensorRT?

Batch normalization can be implemented using a sequence of IElementWiseLayer in TensorRT. More specifically:

adjustedScale = scale / sqrt(variance + epsilon)
batchNorm = (input + bias - (adjustedScale * mean)) * adjustedScale
Why does my network run slower when using DLA than without DLA?

DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments in TensorRT 11.4.0. Refer to Working with DLA for the platforms and packages that support DLA.

DLA was designed to maximize energy efficiency. Depending on the features supported by DLA and the features supported by the GPU, either implementation can be more performant. Your chosen implementation depends on your latency or throughput requirements and power budget. Since all DLA engines are independent of the GPU and each other, you could also use both implementations to increase the throughput of your network further.

Does TensorRT support INT4 quantization or INT16 quantization?

TensorRT supports INT4 quantization for GEMM weight-only quantization. TensorRT does not support INT16 quantization.

What should I do if the ONNX parser does not support a layer my network requires?

The TensorRT ONNX parser is an open-source project. You can add support for custom operators through TensorRT plugins.

Can I use multiple TensorRT builders to compile on different targets?

Warning

TensorRT assumes that all resources for the device or Green Context it is building on are available for optimization purposes. Concurrent use of multiple TensorRT builders (such as multiple trtexec instances) to compile on different targets (DLA0, DLA1, and GPU) can oversubscribe system resources causing undefined behavior (meaning, inefficient plans, builder failure, or system instability).

Using trtexec with the --saveEngine argument, it is recommended to compile for different targets (DLA and GPU) separately and save their plan files. Such plan files can then be reused for loading (using trtexec with the --loadEngine argument) and submitting multiple inference jobs on the respective targets (DLA0, DLA1, and GPU). This two-step process alleviates over-subscription of system resources during the build phase while also allowing execution of the plan file to proceed without interference by the builder.

On a shared CI host, enforce this separation with a per-GPU scheduler, mutex, or semaphore so only one builder profiles a given GPU at a time. Separate processes alone do not prevent resource contention.

DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments in TensorRT 11.4.0. The following DLA0/DLA1 guidance applies only where DLA is supported for your package.

Which layers are accelerated by Tensor Cores?

Most math-bound operations will be accelerated with tensor cores: convolution, deconvolution, fully connected, and matrix multiply. In some cases, particularly for small channel counts or small group sizes, another implementation can be faster and be selected instead of a tensor core implementation.

Why are reformatting layers observed, although there is no warning message that no implementation obeys reformatting-free rules?

Reformat-free network I/O does not mean reformatting layers are not inserted into the entire network. Only the input and output network tensors can be configured not to allow reformatting layers; in other words, TensorRT can insert reformatting layers for internal tensors to improve performance.