Prerequisites#
Before installing TensorRT, ensure your system meets the following requirements. This page is organized by category to help you quickly find the information you need.
Quick Checklist:
NVIDIA GPU (Turing architecture or later)
CUDA Toolkit installed (refer to Which CUDA should I install? below)
Appropriate GPU drivers: r580+ on Linux and Windows for the CUDA 13.x packages, which is every Debian, RPM, tar, and zip package in this release. The lower r535+ on Linux and r537+ on Windows minimum applies only to CUDA 12.x pip wheels
Python 3.10-3.14 supported. Python 3.8 and 3.9 bindings are deprecated, scheduled for removal, and not supported by the samples
On Linux, collect the installed GPU, driver, toolkit, and Python versions with:
nvidia-smi --query-gpu=name,driver_version --format=csv
nvcc --version
python3 --version
nvidia-smi reports the installed driver and GPU. nvcc --version reports
the installed CUDA Toolkit; the CUDA version displayed in the standard
nvidia-smi header is the maximum version supported by the driver, not
necessarily the installed toolkit. On Windows, run nvidia-smi,
nvcc --version, and py --version in PowerShell.
Important
Which CUDA should I install for TensorRT 11.4.0?
Debian / RPM / tar / zip packages: Install CUDA Toolkit 13.4 update 1 and download the TensorRT package whose filename contains
cuda-13.4. This matches the path used by Build Your First Engine.Pip wheels: Use a CUDA 12.x or 13.x driver/toolkit that matches the wheel suffix (
tensorrt/tensorrt-cu13for CUDA 13, ortensorrt-cu12for CUDA 12). Pip does not includetrtexec.Containers: The NGC image pins the CUDA toolkit for you.
Before You Begin#
Review Release Information
Before installation, familiarize yourself with the NVIDIA TensorRT Release Notes to understand:
New features in this release
Known issues and limitations
Platform-specific considerations
Compatibility changes
Choose Your API
TensorRT provides both C++ and Python APIs:
C++ API - Full functionality, no Python dependency
Python API - Convenient for rapid prototyping and integration
Both - Most users install both (default)
The installation instructions assume you want both APIs. For C++ only installation, skip Python-specific packages.
Required Software#
NVIDIA CUDA Toolkit
TensorRT requires the NVIDIA CUDA Toolkit. If not already installed, refer to the NVIDIA CUDA Installation Guide.
Note
TensorRT 11.4.0 Debian, RPM, tar, and zip packages are published for CUDA 13.4 update 1. When downloading from the TensorRT download page, select the cuda-13.4 variant to match the package filename. Prefer CUDA Toolkit 13.4 update 1 for those packages; use CUDA 12.x only when you intentionally install matching pip wheels (tensorrt-cu12) or an older toolkit path documented in the support matrix.
Supported CUDA Versions:
Driver Requirements:
The driver minimum follows the CUDA major version you install, not the operating system.
CUDA 13.x (every Debian, RPM, tar, and zip package in TensorRT 11.4.0, and the default
tensorrtandtensorrt-cu13pip wheels): NVIDIA driver r580 or later on both Linux and Windows.CUDA 12.x (the
tensorrt-cu12pip wheels only): NVIDIA driver r535 or later on Linux, r537 or later on Windows.
For more information, refer to the TensorRT Support Matrix.
Optional Dependencies#
The following libraries are optional and only needed for specific use cases:
cuBLAS (Optional)
cuBLAS is optional and only used for a few specific layers.
When needed: If your model requires cuBLAS-accelerated layers
Installation: Refer to the NVIDIA cuBLAS website
CUDA-Python (Optional)
CUDA-Python enables direct CUDA kernel calls from Python.
When needed: If you use TensorRT Python API with custom CUDA operations
Installation: Refer to the NVIDIA CUDA-Python documentation
NCCL (Optional)
Required only for the multi-device inference feature.
When needed: When using
IDistCollectiveLayer(SM 80+ / NVIDIA Ampere architecture and later) or multi-device attention throughIAttention::setNbRanks(SM 100+ / NVIDIA Blackwell architecture and later)Installation: Refer to the NVIDIA NCCL Installation Guide. The Deep Learning Framework Containers include a compatible NCCL build.
B300 platforms: When using multi-device inference on NVIDIA B300, use NCCL 2.30.x or later to avoid long cold-initialization latency on the first
ncclCommInitRankcall. Refer to the TensorRT 11.0.0 release notes (Known Issues) and TensorRT 11.1.0 release notes (Fixed Issues) for details.
For more information, refer to the TensorRT Support Matrix.
Framework and Model Support#
PyTorch Integration
If you plan to use TensorRT with PyTorch:
Tested with: PyTorch >= 2.0
Compatibility: May work with older versions
Use case: Examples and integration samples
ONNX Model Support
The ONNX-TensorRT parser supports:
ONNX version: 1.20.0
Opset support: Up to opset 25
Backward compatibility: Official support is provided for opset 9 and above
For more information, refer to the TensorRT Support Matrix and the ONNX Opset Guide.
TensorRT Installation Modes#
TensorRT offers three runtime package families:
Full:
tensorrtfor pip or thetensorrtandtensorrt-libspackage family for Debian/RPM. Use Full to build and run engines.Lean:
tensorrt-leanfor pip orlibnvinfer-leanfor Debian/RPM. Use Lean to run compatible version-compatible engines.Dispatch:
tensorrt-dispatchfor pip orlibnvinfer-dispatchfor Debian/RPM. Use Dispatch to select a compatible Lean runtime for a version-compatible engine.
Lean and Dispatch have smaller footprints than Full. Exact installed sizes vary by platform, CUDA variant, and package revision; use your package manager’s metadata for byte counts. The Installation Method Comparison summarizes install paths, and Understanding TensorRT Runtime Options describes when to choose each runtime.
For a development-to-production workflow, use Full Installation during development, serialize optimized engines to plan files, then deploy with Lean or Dispatch in production.
Next Steps#
After verifying prerequisites:
Proceed to Installing TensorRT to choose your installation method
Follow the step-by-step installation instructions for your platform
After verification, follow Build Your First Engine to confirm your install with a 10-minute end-to-end build