DLA Standalone Mode#

Important

DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments in TensorRT 11.4.0. Refer to Working with DLA for the platforms and packages that support DLA.

If you need to run inference outside of TensorRT, you can use EngineCapability::kDLA_STANDALONE to generate a DLA loadable instead of a TensorRT engine. This loadable can then be used with the cuDLA API.

Building A DLA Loadable#

  1. Set the default device type and engine capability to DLA standalone mode.

    builderConfig->setDefaultDeviceType(DeviceType::kDLA);
    builderConfig->setEngineCapability(EngineCapability::kDLA_STANDALONE);
    
    builder_config.default_device_type = trt.DeviceType.DLA
    builder_config.engine_capability = trt.EngineCapability.DLA_STANDALONE
    
  2. Specify the desired precision (FP16, INT8, or mixed) with network tensor types rather than builder flags. In TensorRT 11.0 and later, all networks are strongly typed and NetworkDefinitionCreationFlag::kSTRONGLY_TYPED is deprecated and ignored. Tensor types in the network or a pre-quantized ONNX import drive precision. Refer to Strongly Typed Networks for DLA for DLA-specific guidance.

  3. DLA standalone mode disallows reformatting; therefore, BuilderFlag::kDIRECT_IO needs to be set.

    builderConfig->setFlag(BuilderFlag::kDIRECT_IO);
    
    builder_config.set_flag(trt.BuilderFlag.DIRECT_IO)
    
  4. Set the allowed formats for I/O tensors to one or more of those that are DLA-supported.

  5. Build as normal.

Using trtexec To Generate A DLA Loadable#

The trtexec tool can generate a DLA loadable instead of a TensorRT engine. Specifying both --useDLACore and --safe parameters sets the builder capability to EngineCapability::kDLA_STANDALONE. Specifying --inputIOFormats and --outputIOFormats restricts I/O data type and memory layout. The DLA loadable is saved into a file by specifying --saveEngine parameter.

For example, to generate an FP16 DLA loadable for an ONNX model using trtexec, run:

./trtexec --onnx=model.onnx --saveEngine=model_loadable.bin --useDLACore=0 --inputIOFormats=fp16:chw16 --outputIOFormats=fp16:chw16 --skipInference --safe

TensorRT 11.x removed the per-precision trtexec flags, so the compute precision comes from the network tensor types rather than from --fp16. Supply a strongly typed FP16 ONNX model, which you can produce with ModelOpt AutoCast. The --inputIOFormats and --outputIOFormats values above constrain the I/O data type and layout, which is a separate control from compute precision. Refer to Strongly Typed Networks for DLA.