Installation Guide Overview#
This guide provides complete instructions for installing, upgrading, and uninstalling TensorRT on supported platforms. Whether you are setting up TensorRT for the first time or upgrading an existing installation, this guide will walk you through the process.
About TensorRT#
NVIDIA TensorRT is a high-performance deep learning inference SDK. It includes:
Inference Optimizer - Optimizes trained models for efficient GPU execution
Runtime Engine - Executes optimized models with minimal latency
C++ and Python APIs - Build and deploy inference applications
ONNX Parser - Import models from popular training frameworks
Mixed-Precision Support - FP32, TF32, FP16, BF16, FP8, FP4, INT8, and INT4 inference
TensorRT takes a trained network and produces a highly optimized runtime engine. It applies graph optimizations, layer fusions, and kernel auto-tuning to maximize performance on NVIDIA GPUs from Turing architecture onwards.
In this guide
Prerequisites: system requirements, supported platforms, and required dependencies
Installing TensorRT: install with pip, Debian or RPM packages, tar or zip files, or a container image
Upgrading TensorRT: upgrade from an earlier TensorRT version and manage compatibility
Uninstalling TensorRT: remove TensorRT for each installation method