Strongly Typed Networks for DLA#

Important

DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments in TensorRT 11.4.0. Refer to Working with DLA for the platforms and packages that support DLA.

All networks are strongly typed, including DLA networks. Because DLA on NVIDIA Orin is delivered through the SBSA package, it is not available in TensorRT 11.4.0.

The following builder flags are not valid for DLA builds:

  • BuilderFlag::kFP16

  • BuilderFlag::kINT8

  • BuilderFlag::kPREFER_PRECISION_CONSTRAINTS

  • BuilderFlag::kOBEY_PRECISION_CONSTRAINTS

Network tensor types, including types imported from an ONNX model, determine precision. An ONNX model that was never converted to FP16 keeps its original tensor types, so converting the model is what enables FP16 compute on DLA.

To migrate an existing DLA model, use the NVIDIA TensorRT Model Optimizer ONNX AutoCast tool to convert it to a strongly typed, FP16 representation:

python3 -m modelopt.onnx.autocast \
    --low_precision_type fp16 \
    --onnx_path model.onnx \
    --output_path model_fp16.onnx \
    --opset 20

If some operators do not convert to FP16 as expected, use --data_max to set the maximum absolute activation value allowed before an operator is kept in FP32. AutoCast also converts the network I/O tensor types to FP16 by default. If your deployment requires FP32 I/O, preserve the original I/O types during conversion using the keep_io_types option shown in Python Migration Patterns for TensorRT 11.x.

For general strong-typing behavior and migration guidance, refer to Strongly Typed Networks and Migrating from TensorRT 10.x to 11.x.

FP16 Compute#

For FP16 inference, DLA computes in FP16. TensorRT infers this precision from the network tensor types; no precision builder flag is required. Ensure that the network tensors intended for DLA are typed as FP16 before building.

When parsing from ONNX, kADJUST_FOR_DLA can modify layers and tensor shapes to make them more amenable to DLA execution:

parser.set_flag(trt.OnnxParserFlag.ADJUST_FOR_DLA)

Add explicit cast operations for FP32-to-FP16 conversions at DLA subgraph boundaries. These boundary operations require GPU fallback. Layout transformations, such as NCHW16 to NCHW, do not need to be included in the ONNX model.

Quantization#

TensorRT 11 does not support implicit quantization. For workflows that transition between FP16 and quantized tensors, add explicit IQuantizeLayer (FP16 to UINT8) and IDequantizeLayer (UINT8 to FP16) layers. In ONNX models, use QuantizeLinear and DequantizeLinear nodes.

To enable UINT8 and asymmetric quantization while parsing an ONNX model, set kENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLA:

parser.set_flag(
    trt.OnnxParserFlag.ENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLA
)

The flag is not required for symmetric INT8 quantization. DLA supports asymmetric quantization for UINT8 (nonzero zero-point values).

For validated Windows on ARM shapes, refer to IQuantizeLayer and IDequantizeLayer in Layer Details.

Models whose tensors are specified natively as INT8 do not require explicit quantization nodes throughout the network. Explicit quantization remains the recommended way to represent transitions between FP16 and INT8.

DLA Data Types#

DLA supports FP16 and INT8 compute. It does not support FP8, BF16, or other reduced-precision compute types. NVIDIA TensorRT Model Optimizer AutoCast and quantization tools can prepare an ONNX model before parsing. DLA supports the resulting QuantizeLinear and DequantizeLinear nodes when kENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLA is set for UINT8 or asymmetric quantization.