Working with DLA#

Important

Starting with TensorRT 11.4.0, DLA support depends on the platform and package:

  • Supported: Enterprise TensorRT 11.4.0 on platforms including Windows on ARM, where DLA support is in beta.

  • Not supported: Linux SBSA, NVIDIA JetPack, and NVIDIA DriveOS deployments. DLA was also unavailable in TensorRT 11.0 through 11.3, including TensorRT 11.3.1 for DriveOS. If DLA is required on any of these platforms, remain on TensorRT 10.7.

DLA support on Windows on ARM is in beta and covers a restricted set of layers. The supported-layer set and restrictions differ from Jetson and DriveOS: notably, there is no GPU fallback, so every layer must be DLA-capable, and batch size is 1. DLA engines built with TensorRT 11.4.0 must be rebuilt once DLA support reaches general availability, so keep your source models. Refer to DLA Supported Layers for Windows on ARM for that platform.

NVIDIA DLA (Deep Learning Accelerator) is a fixed-function accelerator engine targeted for deep learning operations. It is designed to fully hardware accelerate convolutional neural networks. DLA supports various layers, such as convolution, deconvolution, fully connected, activation, pooling, and batch normalization. Refer to the DLA Supported Layers and Restrictions section for Jetson and DriveOS DLA support.

DLA is useful for offloading CNN processing from the GPU and is significantly more power-efficient for these workloads. In addition, it can provide an independent execution pipeline in cases where redundancy is important, such as mission-critical or safety applications.

For more information, refer to the DLA Developer page and the DLA tutorial Getting Started with the Deep Learning Accelerator on NVIDIA Jetson Orin.

When building a model for DLA, the TensorRT builder parses the network and calls the DLA compiler to compile the network into a DLA loadable. Refer to Using trtexec to learn how to build and run networks on DLA.

DLA Software Stack: Build and Runtime Phases

In this guide