DLA Supported Layers for Windows on ARM#

This page describes the TensorRT layers that DLA supports on Windows on ARM platforms.

Important

The Windows on ARM DLA supported-operation set and restrictions differ from Jetson and DriveOS DLA. Do not apply the limits on this page to Jetson or DriveOS builds. For those platforms, refer to DLA Supported Layers and Restrictions.

Note

DLA support on Windows on ARM is in beta, and the layers on this page are a restricted set. DLA engines built with TensorRT 11.4.0 must be rebuilt once DLA support reaches general availability. Refer to the TensorRT 11.4.0 Release Notes.

General Restrictions#

The following restrictions apply to all layers:

  • TensorRT accepts tensors of any rank. Spatial dimensions (C, H, W) are the last three dimensions of each tensor. Individual layers can require a specific rank (often 4D). Refer to Layer Details.

  • Default DLA compute precision is FP16. Some layers also support FP32, as noted in the support table and layer details.

  • Asymmetric quantization through IQuantizeLayer / IDequantizeLayer is supported as a quantization format only. Compute precision remains FP16. Pure INT8 compute is not supported on this platform.

  • GPU fallback is not supported. Every layer in the network must be DLA-capable. A network that contains any unsupported layer fails to build.

  • Dynamic shapes are not supported. All dimensions must be fixed at build time.

  • Batch size must be 1.

  • Mixed-precision networks (FP32 network I/O with FP16 DLA compute) are supported using the FP32 and FP16 boundary formats listed in I/O Formats on DLA.

This page does not provide representative accuracy or performance results. Before deployment, compare DLA output accuracy and latency with the GPU baseline for your network. The 11.4.0 release notes also document an intermittent DLA accuracy issue that must be included in that validation.

Layer Support#

The following table summarizes the layers supported on Windows on ARM DLA, including tensor data types and the main constraints for each layer. Shapes in the following table are CHW. Leading dimensions can vary. For the full constraint set, refer to Layer Details.

Table 16 Windows on ARM DLA Layer Support Summary#

Layer

Tensor data types

Key constraints

Activation (IActivationLayer)

FP16, FP32

kRELU, kSIGMOID, kTANH, kLEAKY_RELU

Concatenation (IConcatenationLayer)

FP16, FP32

Exact axis and input-count tuples listed below; up to 40 inputs validated

Constant (IConstantLayer)

N/A

FP32, INT32, INT8, or UINT8 values

Convolution (IConvolutionLayer)

FP16, FP32

2D; regular and depthwise; dilation [1,1] only; exact validated strides listed below, including one [16,16] patch-embedding case

Dequantize (IDequantizeLayer)

UINT8→FP32 or FP16 at I/O boundaries

Axis=1; asymmetric; FP32 scale, UINT8 zero-point

ElementWise (IElementWiseLayer)

FP16, FP32

kSUM, kSUB, kPROD, kDIV; comparison subgraph patterns listed below; broadcasting on size-1 dimensions

Normalization (INormalizationLayer)

FP16

Layer normalization only; axes over last dimension or last two dimensions; constant scale and bias

Pooling (IPoolingLayer)

FP16, FP32

kMAX and kAVERAGE; exact validated window and stride tuples listed below, including a 7×7 window and strides up to 2×2

Quantize (IQuantizeLayer)

FP16→UINT8

Axis=1; asymmetric; FP32 scale, UINT8 zero-point

Shuffle (IShuffleLayer)

FP16, FP32

Reshape, transpose, depth-to-space, and space-to-depth patterns are identified, but no exact validated tuples are available on this page

Note

The configurations listed in Layer Details are specific validated tuples of operation type, data type, shape, and attributes from the DLA Compiler test suite. They are not a Cartesian product; only the exact combinations shown have been validated. Broader support can exist but has not been confirmed. If a required configuration is not listed, attempt an engine build and report a failure through Reporting TensorRT Issues.

Refer to each layer’s detailed restrictions in Layer Details for the full set of constraints.

Layer Details#

Activation
  • Tensor data types: FP16 or FP32 for input tensors; FP16 for outputs. Compute precision is fixed at FP16.

  • Supported activation types (ActivationType): kRELU, kSIGMOID, kTANH, kLEAKY_RELU.

  • Validated (ActivationType, [C,H,W]) pairs:

    • kRELU:

      • [32,7,7]

      • [32,14,14]

      • [32,32,40]

      • [32,32,80]

      • [32,32,160]

      • [32,32,320]

      • [32,32,4000]

      • [32,112,112]

      • [48,7,7]

      • [48,14,14]

      • [64,7,7]

      • [64,28,28]

      • [64,48,72]

      • [64,96,144]

      • [64,112,112]

      • [64,192,288]

      • [64,384,576]

      • [96,7,7]

      • [128,16,20]

      • [128,16,40]

      • [128,16,80]

      • [128,16,160]

      • [128,16,2000]

      • [128,56,56]

      • [128,96,144]

      • [16,64,80]

      • [16,64,160]

      • [16,64,320]

      • [16,64,640]

      • [16,64,8000]

      • [256,8,20]

      • [256,8,40]

      • [256,8,80]

      • [256,8,160]

      • [256,8,1000]

      • [256,48,72]

      • [512,4,20]

      • [512,4,40]

      • [512,4,80]

      • [512,4,160]

      • [512,4,500]

      • [512,24,36]

      • [500,1,32]

    • kSIGMOID:

      • [1,48,72]

      • [1,96,144]

      • [1,128,224]

      • [1,192,288]

      • [1,256,448]

      • [1,300,8]

      • [1,300,256]

      • [1,300,2048]

      • [1,3600,256]

      • [8,48,72]

      • [8,96,144]

      • [8,192,288]

      • [12,77,2]

      • [12,256,2]

      • [12,512,2]

      • [16,225,768]

      • [128,1,1]

      • [128,60,60]

      • [256,60,60]

    • kTANH:

      • [1,1,768]

      • [10,64,128]

      • [128,1,1]

      • [3,512,512]

    • kLEAKY_RELU (alpha=0.3):

      • [12,256,256]

      • [32,256,256]

      • [64,128,128]

Concatenation
  • Tensor data types: FP16 or FP32.

  • The displayed input shapes omit the fixed leading batch dimension. Axis values are full NCHW tensor axes, so axis=0 is the batch axis. The validated axis=0 cases use one input and therefore preserve batch size 1.

  • Up to 40 inputs have been validated.

  • Validated configurations (input [C,H,W] per tensor × num_inputs, axis):

    • axis=0:

      • 1× [1,1,1024]

      • 1× [1,1,32128]

      • 1× [1,64,1024]

    • axis=1:

      • 2× [128,28,28]

      • 2× [128,7,7] + [32,7,7]

      • 2× [160,7,7] + [64,7,7]

      • 2× [128,14,14] + [384,14,14]

      • 2× [32,7,7] + [64,7,7]

      • 2× [32,7,7] + [96,7,7]

      • 3× [192,60,60]

      • 5× [128,60,60]

    • axis=2:

      • 2× [128,1,160]

      • 2× [128,1,20]

      • 2× [128,1,40]

      • 2× [128,1,80]

    • axis=3:

      • 2× [1,1,32]

      • 2× [1,64,32]

      • 2× [1,300,2]

      • 2× [1,3600,2]

      • 2× [12,256,1]

      • 2× [12,512,1]

      • 2× [12,77,1]

      • 2× [64,1,32]

      • 2× [300,64,1]

      • 4× [1,300,128]

      • 20× [128,1,1]

      • 40× [128,1,1]

Constant
  • Produces a constant tensor consumed by other DLA layers.

  • Supported value types: FP32, INT32, INT8, UINT8.

    • FP32: general-purpose values, including scale tensors, bias tensors, and convolution weights.

    • INT8: valid only as quantized convolution weight kernels consumed by IConvolutionLayer.

    • UINT8: valid only as the data input to IDequantizeLayer, or as the zero-point input to IQuantizeLayer / IDequantizeLayer.

    • INT32: valid only as index or shape tensors consumed by layers that accept integer types (for example, ISliceLayer, IGatherLayer).

  • Validated [C,H,W] shapes, grouped by value type:

    • FP32 (4):

      • [1,1,128]

      • [12,64,512]

      • [3,1,1]

      • [8,300,32]

    • INT32 (2):

      • [1,1,1]

      • [12,64,64]

    • INT8 (27):

      • [1,3,3]

      • [3,3,3]

      • [4,3,3]

      • [8,7,7]

      • [16,3,3]

      • [24,1,1]

      • [24,3,3]

      • [32,1,1]

      • [32,3,3]

      • [48,3,3]

      • [60,3,3]

      • [64,1,1]

      • [64,3,3]

      • [64,4,4]

      • [64,7,7]

      • [96,3,3]

      • [128,1,1]

      • [128,3,3]

      • [128,4,4]

      • [132,3,3]

      • [160,3,3]

      • [224,3,3]

      • [256,1,1]

      • [256,3,3]

      • [512,1,1]

      • [512,3,3]

      • [1024,1,1]

    • UINT8 (20):

      • [1,3,3]

      • [3,7,7]

      • [3,16,16]

      • [64,1,1]

      • [64,3,3]

      • [96,1,1]

      • [128,1,1]

      • [128,3,3]

      • [192,1,1]

      • [256,1,1]

      • [256,3,3]

      • [512,1,1]

      • [512,3,3]

      • [576,1,1]

      • [640,1,1]

      • [768,1,1]

      • [1024,1,1]

      • [2048,1,1]

      • [3072,1,1]

      • [4096,1,1]

Convolution
  • Tensor data types: FP16 or FP32 for the input feature map; FP32 for weights and bias.

  • 2D convolutions only.

  • Regular convolutions (num_groups=1) and depthwise convolutions (num_groups equals the number of input channels) are supported.

  • Dilation must be [1,1] (no dilated convolutions).

  • Stride values up to [16,16] have been validated (for patch-embedding kernels).

  • Bias mode: CHANNEL only.

  • Validated configurations, grouped by input [C,H,W] (kernel [C/g,R,S], stride; * denotes depthwise):

    • [3,60,8000]: [3,3,3] s=[1,1]

    • [3,224,224]: [3,3,3] s=[1,1]; [3,3,3] s=[2,2]

    • [3,768,1152]: [3,7,7] s=[2,2]

    • [3,960,960]: [3,16,16] s=[16,16]

    • [4,128,224]: [4,3,3] s=[2,2]

    • [16,16,2000]: [16,3,3] s=[1,1]

    • [16,32,4000]: [16,3,3] s=[1,1]

    • [24,64,112]: [24,3,3] s=[1,1]; [24,3,3] s=[2,2]

    • [24,128,224]: [24,3,3] s=[1,1]

    • [32,8,1000]: [32,3,3] s=[1,1]

    • [32,14,14]: [32,3,3] s=[1,1]

    • [32,16,2000]: [32,1,1] s=[1,1]

    • [32,112,112]: *[1,3,3] s=[1,1]; [32,1,1] s=[1,1]; [32,3,3] s=[1,1]; [32,3,3] s=[2,2]

    • [32,224,224]: [32,3,3] s=[2,2]

    • [48,7,7]: [48,3,3] s=[1,1]

    • [48,14,14]: [48,3,3] s=[1,1]; [48,3,3] s=[2,2]

    • [60,32,56]: [60,3,3] s=[1,1]; [60,3,3] s=[2,2]

    • [64,4,500]: [64,3,3] s=[1,1]

    • [64,7,7]: [64,3,3] s=[1,1]

    • [64,28,28]: [64,3,3] s=[1,1]; [64,3,3] s=[2,2]

    • [64,48,72]: [64,1,1] s=[1,1]; [64,3,3] s=[1,1]

    • [64,56,56]: [64,1,1] s=[1,1]; [64,3,3] s=[1,1]; [64,3,3] s=[2,2]

    • [64,96,144]: [64,1,1] s=[1,1]; [64,3,3] s=[1,1]

    • [64,112,112]: *[1,3,3] s=[2,2]; [64,1,1] s=[1,1]

    • [64,192,288]: [64,1,1] s=[1,1]; [64,1,1] s=[2,2]; [64,3,3] s=[1,1]; [64,3,3] s=[2,2]

    • [96,7,7]: [96,3,3] s=[1,1]

    • [96,16,28]: [96,3,3] s=[1,1]; [96,3,3] s=[2,2]

    • [128,7,7]: [128,3,3] s=[1,1]

    • [128,8,1000]: [128,1,1] s=[1,1]

    • [128,16,2000]: [128,1,1] s=[1,1]

    • [128,28,28]: [128,1,1] s=[1,1]; [128,3,3] s=[1,1]; [128,3,3] s=[2,2]

    • [128,56,56]: *[1,3,3] s=[1,1]; *[1,3,3] s=[2,2]; [128,1,1] s=[1,1]; [128,3,3] s=[2,2]

    • [128,60,60]: [128,3,3] s=[1,1]

    • [128,96,144]: [128,1,1] s=[1,1]; [128,1,1] s=[2,2]; [128,3,3] s=[1,1]; [128,3,3] s=[2,2]

    • [132,8,14]: [132,3,3] s=[1,1]

    • [160,7,7]: [160,3,3] s=[1,1]

    • [224,7,7]: [224,3,3] s=[1,1]

    • [256,4,500]: [256,1,1] s=[1,1]

    • [256,7,7]: [256,3,3] s=[1,1]

    • [256,8,1000]: [256,1,1] s=[1,1]

    • [256,14,14]: [256,3,3] s=[1,1]; [256,3,3] s=[2,2]

    • [256,28,28]: *[1,3,3] s=[1,1]; [256,1,1] s=[1,1]; [256,3,3] s=[1,1]; [256,3,3] s=[2,2]

    • [256,48,72]: [256,1,1] s=[1,1]; [256,1,1] s=[2,2]; [256,3,3] s=[1,1]; [256,3,3] s=[2,2]

    • [512,4,500]: [512,1,1] s=[1,1]

    • [512,7,7]: [512,1,1] s=[1,1]; [512,3,3] s=[1,1]

    • [512,14,14]: *[1,3,3] s=[1,1]; [512,1,1] s=[1,1]; [512,3,3] s=[2,2]

    • [512,24,36]: [512,1,1] s=[1,1]; [512,3,3] s=[1,1]

    • [576,60,60]: [576,1,1] s=[1,1]

    • [640,60,60]: [640,1,1] s=[1,1]

    • [768,1,1]: [768,1,1] s=[1,1]

    • [1024,1,1]: [1024,1,1] s=[1,1]

    • [1024,7,7]: *[1,3,3] s=[1,1]

    • [1024,64,1]: [1024,1,1] s=[1,1]

Dequantize
  • Input: UINT8. Output: FP32 (or FP16 at I/O boundaries).

  • Axis=1; asymmetric quantization only (zero-point ≠ 0 permitted).

  • Scale: FP32 scalar per channel. Zero-point: UINT8 scalar per channel.

  • Validated input [C,H,W] shapes:

    • [1,1,112]

    • [1,1,500]

    • [1,1,512]

    • [1,1,1000]

    • [1,1,2000]

    • [1,1,4000]

    • [1,1,8000]

    • [1,16,4]

    • [1,48,72]

    • [1,96,144]

    • [1,128,224]

    • [1,192,288]

    • [1,256,448]

    • [3,224,224]

    • [3,60,8000]

    • [3,768,1152]

    • [4,128,224]

    • [8,48,72]

    • [8,96,144]

    • [8,192,288]

    • [16,4,500]

    • [16,16,2000]

    • [16,32,4000]

    • [16,64,8000]

    • [24,32,56]

    • [24,64,112]

    • [24,128,224]

    • [32,7,7]

    • [32,8,1000]

    • [32,14,14]

    • [32,16,2000]

    • [32,32,4000]

    • [32,112,112]

    • [32,224,224]

    • [48,7,7]

    • [48,14,14]

    • [60,16,28]

    • [60,32,56]

    • [64,7,7]

    • [64,4,500]

    • [64,24,36]

    • [64,28,28]

    • [64,48,72]

    • [64,56,56]

    • [64,96,144]

    • [64,112,112]

    • [64,192,288]

    • [64,384,576]

    • [70,28,28]

    • [96,7,7]

    • [96,8,14]

    • [96,16,28]

    • [128,8,1000]

    • [128,14,14]

    • [128,16,2000]

    • [128,28,28]

    • [128,56,56]

    • [128,96,144]

    • [132,8,14]

    • [256,4,500]

    • [256,7,7]

    • [256,8,1000]

    • [256,14,14]

    • [256,28,28]

    • [256,48,72]

    • [384,14,14]

    • [500,1,2]

    • [500,1,10]

    • [500,1,32]

    • [500,1,64]

    • [512,1,1]

    • [512,4,500]

    • [512,7,7]

    • [512,14,14]

    • [512,24,36]

ElementWise
  • Tensor data types: FP16 or FP32.

  • Supported operations (ElementWiseOperation): kSUM, kSUB, kPROD, kDIV.

  • Comparison operations are supported only in these complete subgraph patterns:

    • kEQUAL, kGREATER, or kLESS immediately followed by a cast to FP16 or INT8.

    • kGREATER or kLESS followed by kNOT and then a cast to FP16 or INT8.

    The INT8 cast in these patterns is a supported intermediate conversion; it does not add general pure-INT8 compute support.

  • Broadcasting is supported along dimensions whose size is 1.

  • Validated configurations (input1 [C,H,W] OP input2 [C,H,W]):

    • kDIV:

      • [1,1,1024] OP [1,1,1]

      • [1,1,768] OP [1,1,1]

      • [1,1,768] OP [1,1,768]

      • [1,64,1024] OP [1,64,1]

      • [256,60,60] OP [1,60,60]

    • kSUB:

      • [1,300,1] OP [1,300,1]

      • [256,60,60] OP [1,60,60]

    • kPROD:

      • [1,1,100] OP [1,1,100]

      • [1,1,1000] OP [256,8,1000]

      • [1,1,1000] OP [32,8,1000]

      • [1,1,1024] OP [1,1,1024]

      • [1,1,2000] OP [128,16,2000]

      • [1,1,2000] OP [16,16,2000]

      • [1,1,4000] OP [32,32,4000]

      • [1,1,500] OP [16,4,500]

      • [1,1,500] OP [512,4,500]

      • [1,1,500] OP [64,4,500]

      • [1,1,768] OP [1,1,768]

      • [1,1,8000] OP [16,64,8000]

      • [1,300,2048] OP [1,300,2048]

      • [1,300,256] OP [1,300,256]

      • [1,3600,256] OP [1,3600,256]

      • [1,64,1024] OP [1,64,1024]

      • [1,64,768] OP [1,64,1]

      • [12,256,1] OP [12,256,1]

      • [12,512,1] OP [12,512,1]

      • [12,77,1] OP [12,77,1]

      • [128,60,60] OP [128,60,60]

      • [16,225,768] OP [16,225,768]

      • [16,300,4] OP [1,300,4]

      • [256,60,60] OP [256,60,60]

      • [300,64,2] OP [300,1,2]

    • kSUM:

      • [1,1,64] OP [1,1,64]

      • [1,1,768] OP [1,1,768]

      • [1,1,1024] OP [1,1,1024]

      • [1,49,1024] OP [1,49,1024]

      • [1,64,64] OP [1,1,64]

      • [1,64,64] OP [1,64,64]

      • [1,64,768] OP [1,64,768]

      • [1,64,1024] OP [1,64,1024]

      • [1,196,512] OP [1,196,512]

      • [1,256,768] OP [1,256,768]

      • [1,300,1] OP [1,300,1]

      • [1,300,2] OP [1,300,2]

      • [1,512,768] OP [1,512,768]

      • [1,784,256] OP [1,784,256]

      • [1,3136,128] OP [1,3136,128]

      • [3,512,512] OP [3,512,512]

      • [12,64,64] OP [12,64,64]

      • [12,77,77] OP [1,1,77]

      • [12,77,77] OP [12,77,77]

      • [12,256,256] OP [1,1,256]

      • [12,256,256] OP [12,256,256]

      • [12,512,512] OP [1,1,512]

      • [12,512,512] OP [12,512,512]

      • [16,225,192] OP [16,225,192]

      • [24,64,112] OP [24,64,112]

      • [60,32,56] OP [60,32,56]

      • [64,1,64] OP [64,1,64]

      • [64,48,72] OP [64,48,72]

      • [64,96,144] OP [64,96,144]

      • [64,112,112] OP [64,112,112]

      • [64,128,128] OP [64,128,128]

      • [64,192,288] OP [64,192,288]

      • [96,16,28] OP [96,16,28]

      • [128,56,56] OP [128,56,56]

      • [128,96,144] OP [128,96,144]

      • [256,16,32] OP [256,16,32]

      • [256,28,28] OP [256,28,28]

      • [256,48,72] OP [256,48,72]

      • [300,1,2] OP [300,64,2]

      • [32,256,256] OP [32,256,256]

      • [512,7,7] OP [512,7,7]

      • [512,14,14] OP [512,14,14]

      • [512,24,36] OP [512,24,36]

Normalization
  • Tensor data types: FP16.

  • Layer normalization only. The axes bitmask must cover only the last dimension (bit 3) or the last two dimensions (bits 2 and 3).

  • Scale and bias tensors must be build-time constants whose size matches the normalized dimension(s).

  • Epsilon values in the range [1e-7, 1e-5] have been validated.

  • Validated [C,H,W] shapes:

    • axes=[3] (last dim, 12 shapes):

      • [1,1,1024]

      • [1,49,1024]

      • [1,64,768]

      • [1,77,768]

      • [1,196,512]

      • [1,256,768]

      • [1,300,256]

      • [1,512,768]

      • [1,784,256]

      • [1,3136,128]

      • [1,3600,256]

      • [16,225,192]

    • axes=[2,3] (last two dims, 3 shapes):

      • [64,64,128]

      • [128,32,64]

      • [256,16,32]

Pooling
  • Tensor data types: FP16 or FP32.

  • Supported types (PoolingType): kMAX, kAVERAGE.

  • Maximum window size: 7×7.

  • Validated configurations (input [C,H,W], window, stride):

    • kAVERAGE:

      • [512,7,7] win=[7,7] stride=[2,2]

    • kMAX:

      • [16,64,80] win=[2,2] stride=[2,2]

      • [16,64,160] win=[2,2] stride=[2,2]

      • [16,64,320] win=[2,2] stride=[2,2]

      • [16,64,640] win=[2,2] stride=[2,2]

      • [16,64,8000] win=[2,2] stride=[2,2]

      • [32,32,40] win=[2,2] stride=[2,2]

      • [32,32,80] win=[2,2] stride=[2,2]

      • [32,32,160] win=[2,2] stride=[2,2]

      • [32,32,320] win=[2,2] stride=[2,2]

      • [32,32,4000] win=[2,2] stride=[2,2]

      • [64,384,576] win=[3,3] stride=[2,2]

      • [128,16,20] win=[2,1] stride=[2,1]

      • [128,16,40] win=[2,1] stride=[2,1]

      • [128,16,80] win=[2,1] stride=[2,1]

      • [128,16,160] win=[2,1] stride=[2,1]

      • [128,16,2000] win=[2,2] stride=[2,2]

      • [256,8,20] win=[2,1] stride=[2,1]

      • [256,8,40] win=[2,1] stride=[2,1]

      • [256,8,80] win=[2,1] stride=[2,1]

      • [256,8,160] win=[2,1] stride=[2,1]

      • [256,8,1000] win=[2,2] stride=[2,2]

Quantize
  • Input: FP16. Output: UINT8.

  • Axis=1; asymmetric quantization only (zero-point ≠ 0 permitted).

  • Scale: FP32 scalar per channel. Zero-point: UINT8 scalar per channel.

  • Validated input [C,H,W] shapes:

    • [1,1,112]

    • [1,1,500]

    • [1,1,512]

    • [1,1,1000]

    • [1,1,2000]

    • [1,1,4000]

    • [1,4,28]

    • [1,16,4]

    • [1,48,72]

    • [1,96,144]

    • [1,128,224]

    • [1,192,288]

    • [1,256,448]

    • [3,224,224]

    • [3,60,8000]

    • [3,768,1152]

    • [4,128,224]

    • [8,48,72]

    • [8,96,144]

    • [8,192,288]

    • [16,4,500]

    • [16,16,2000]

    • [16,32,4000]

    • [16,64,8000]

    • [24,32,56]

    • [24,64,112]

    • [24,128,224]

    • [28,28,70]

    • [32,7,7]

    • [32,8,1000]

    • [32,14,14]

    • [32,16,2000]

    • [32,32,4000]

    • [32,112,112]

    • [32,224,224]

    • [48,7,7]

    • [48,14,14]

    • [60,16,28]

    • [60,32,56]

    • [64,7,7]

    • [64,4,500]

    • [64,24,36]

    • [64,28,28]

    • [64,48,72]

    • [64,56,56]

    • [64,96,144]

    • [64,112,112]

    • [64,192,288]

    • [64,384,576]

    • [70,28,28]

    • [96,7,7]

    • [96,8,14]

    • [96,16,28]

    • [128,8,1000]

    • [128,14,14]

    • [128,16,2000]

    • [128,28,28]

    • [128,56,56]

    • [128,96,144]

    • [128,224,1]

    • [132,8,14]

    • [256,4,500]

    • [256,7,7]

    • [256,8,1000]

    • [256,14,14]

    • [256,28,28]

    • [256,48,72]

    • [256,448,1]

    • [384,14,14]

    • [500,1,2]

    • [500,1,10]

    • [500,1,32]

    • [500,1,64]

    • [512,1,1]

    • [512,4,500]

    • [512,7,7]

    • [512,14,14]

    • [512,24,36]

Shuffle
  • Tensor data types: FP16 or FP32.

  • IShuffleLayer covers reshape, transpose, depth-to-space, and space-to-depth operations. The DLA Compiler test source identifies all four patterns, but this page has no exact validated shape and attribute tuples for them. Treat these patterns as unconfirmed until the target engine builds successfully.

I/O Formats on DLA#

TensorRT supports the following tensor formats for network inputs and outputs. Specify formats with the TensorFormat enum on the network’s input and output bindings.

Table 17 Windows on ARM DLA I/O Formats#

Data type

TensorFormat

Notes

FP16

kCHW16

Supported for both inputs and outputs.

FP32

kLINEAR

Supported for both inputs and outputs.

UINT8

kLINEAR

Supported for outputs only. Requires the network to end with an IQuantizeLayer (FP16→UINT8).

FP16/kCHW16 and FP32/kLINEAR are validated as inputs and outputs in the test suite. UINT8/kLINEAR is validated as an output only.

Note

TensorFormat::kDLA_LINEAR is a DLA-specific planar layout distinct from kLINEAR. Specify kLINEAR at the TensorRT API level. TensorRT maps it to the correct layout internally.

For comparison with Jetson and DriveOS DLA formats, refer to I/O Formats on DLA.

Strong Typing#

Important

All networks are strongly typed, including DLA networks. TensorRT 11.0 removed the weak-typing APIs, so DLA on Windows on ARM requires strongly typed networks. The following builder flags are not valid for DLA builds:

  • BuilderFlag::kFP16

  • BuilderFlag::kINT8

  • BuilderFlag::kPREFER_PRECISION_CONSTRAINTS

  • BuilderFlag::kOBEY_PRECISION_CONSTRAINTS

Use tensor types in the network (or a pre-quantized ONNX import) to set precision. Refer to Strongly Typed Networks for DLA, ONNX Parser Flags for DLA, and Migrating from TensorRT 10.x to 11.x.