Strongly Typed Networks for DLA#
Important
DLA is not supported for Linux SBSA, NVIDIA JetPack, or NVIDIA DriveOS deployments in TensorRT 11.4.0. Refer to Working with DLA for the platforms and packages that support DLA.
All networks are strongly typed, including DLA networks. Because DLA on NVIDIA Orin is delivered through the SBSA package, it is not available in TensorRT 11.4.0.
The following builder flags are not valid for DLA builds:
BuilderFlag::kFP16BuilderFlag::kINT8BuilderFlag::kPREFER_PRECISION_CONSTRAINTSBuilderFlag::kOBEY_PRECISION_CONSTRAINTS
Network tensor types, including types imported from an ONNX model, determine precision. An ONNX model that was never converted to FP16 keeps its original tensor types, so converting the model is what enables FP16 compute on DLA.
To migrate an existing DLA model, use the NVIDIA TensorRT Model Optimizer ONNX AutoCast tool to convert it to a strongly typed, FP16 representation:
python3 -m modelopt.onnx.autocast \
--low_precision_type fp16 \
--onnx_path model.onnx \
--output_path model_fp16.onnx \
--opset 20
If some operators do not convert to FP16 as expected, use --data_max to set
the maximum absolute activation value allowed before an operator is kept in
FP32. AutoCast also converts the network I/O tensor types to FP16 by default.
If your deployment requires FP32 I/O, preserve the original I/O types during
conversion using the keep_io_types option shown in
Python Migration Patterns for TensorRT 11.x.
For general strong-typing behavior and migration guidance, refer to Strongly Typed Networks and Migrating from TensorRT 10.x to 11.x.
FP16 Compute#
For FP16 inference, DLA computes in FP16. TensorRT infers this precision from the network tensor types; no precision builder flag is required. Ensure that the network tensors intended for DLA are typed as FP16 before building.
When parsing from ONNX, kADJUST_FOR_DLA can modify layers and tensor shapes
to make them more amenable to DLA execution:
parser.set_flag(trt.OnnxParserFlag.ADJUST_FOR_DLA)
Add explicit cast operations for FP32-to-FP16 conversions at DLA subgraph boundaries. These boundary operations require GPU fallback. Layout transformations, such as NCHW16 to NCHW, do not need to be included in the ONNX model.
Quantization#
TensorRT 11 does not support implicit quantization. For workflows that
transition between FP16 and quantized tensors, add explicit
IQuantizeLayer (FP16 to UINT8) and IDequantizeLayer (UINT8 to FP16)
layers. In ONNX models, use QuantizeLinear and DequantizeLinear nodes.
To enable UINT8 and asymmetric quantization while parsing an ONNX model, set
kENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLA:
parser.set_flag(
trt.OnnxParserFlag.ENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLA
)
The flag is not required for symmetric INT8 quantization. DLA supports asymmetric quantization for UINT8 (nonzero zero-point values).
For validated Windows on ARM shapes, refer to IQuantizeLayer and
IDequantizeLayer in Layer Details.
Models whose tensors are specified natively as INT8 do not require explicit quantization nodes throughout the network. Explicit quantization remains the recommended way to represent transitions between FP16 and INT8.
DLA Data Types#
DLA supports FP16 and INT8 compute. It does not support FP8, BF16, or other
reduced-precision compute types. NVIDIA TensorRT Model Optimizer AutoCast and
quantization tools can prepare an ONNX model before parsing. DLA supports the
resulting QuantizeLinear and DequantizeLinear nodes when
kENABLE_UINT8_AND_ASYMMETRIC_QUANTIZATION_DLA is set for UINT8 or
asymmetric quantization.