DLA Supported Layers for Windows on ARM#
This page describes the TensorRT layers that DLA supports on Windows on ARM platforms.
Important
The Windows on ARM DLA supported-operation set and restrictions differ from Jetson and DriveOS DLA. Do not apply the limits on this page to Jetson or DriveOS builds. For those platforms, refer to DLA Supported Layers and Restrictions.
Note
DLA support on Windows on ARM is in beta, and the layers on this page are a restricted set. DLA engines built with TensorRT 11.4.0 must be rebuilt once DLA support reaches general availability. Refer to the TensorRT 11.4.0 Release Notes.
General Restrictions#
The following restrictions apply to all layers:
TensorRT accepts tensors of any rank. Spatial dimensions (C, H, W) are the last three dimensions of each tensor. Individual layers can require a specific rank (often 4D). Refer to Layer Details.
Default DLA compute precision is FP16. Some layers also support FP32, as noted in the support table and layer details.
Asymmetric quantization through
IQuantizeLayer/IDequantizeLayeris supported as a quantization format only. Compute precision remains FP16. Pure INT8 compute is not supported on this platform.GPU fallback is not supported. Every layer in the network must be DLA-capable. A network that contains any unsupported layer fails to build.
Dynamic shapes are not supported. All dimensions must be fixed at build time.
Batch size must be
1.Mixed-precision networks (FP32 network I/O with FP16 DLA compute) are supported using the FP32 and FP16 boundary formats listed in I/O Formats on DLA.
This page does not provide representative accuracy or performance results. Before deployment, compare DLA output accuracy and latency with the GPU baseline for your network. The 11.4.0 release notes also document an intermittent DLA accuracy issue that must be included in that validation.
Layer Support#
The following table summarizes the layers supported on Windows on ARM DLA, including tensor data types and the main constraints for each layer. Shapes in the following table are CHW. Leading dimensions can vary. For the full constraint set, refer to Layer Details.
Layer |
Tensor data types |
Key constraints |
|---|---|---|
Activation ( |
FP16, FP32 |
|
Concatenation ( |
FP16, FP32 |
Exact axis and input-count tuples listed below; up to 40 inputs validated |
Constant ( |
N/A |
FP32, INT32, INT8, or UINT8 values |
Convolution ( |
FP16, FP32 |
2D; regular and depthwise; dilation |
Dequantize ( |
UINT8→FP32 or FP16 at I/O boundaries |
Axis=1; asymmetric; FP32 scale, UINT8 zero-point |
ElementWise ( |
FP16, FP32 |
|
Normalization ( |
FP16 |
Layer normalization only; axes over last dimension or last two dimensions; constant scale and bias |
Pooling ( |
FP16, FP32 |
|
Quantize ( |
FP16→UINT8 |
Axis=1; asymmetric; FP32 scale, UINT8 zero-point |
Shuffle ( |
FP16, FP32 |
Reshape, transpose, depth-to-space, and space-to-depth patterns are identified, but no exact validated tuples are available on this page |
Note
The configurations listed in Layer Details are specific validated tuples of operation type, data type, shape, and attributes from the DLA Compiler test suite. They are not a Cartesian product; only the exact combinations shown have been validated. Broader support can exist but has not been confirmed. If a required configuration is not listed, attempt an engine build and report a failure through Reporting TensorRT Issues.
Refer to each layer’s detailed restrictions in Layer Details for the full set of constraints.
Layer Details#
Activation
Tensor data types: FP16 or FP32 for input tensors; FP16 for outputs. Compute precision is fixed at FP16.
Supported activation types (
ActivationType):kRELU,kSIGMOID,kTANH,kLEAKY_RELU.Validated (
ActivationType,[C,H,W]) pairs:kRELU:[32,7,7][32,14,14][32,32,40][32,32,80][32,32,160][32,32,320][32,32,4000][32,112,112][48,7,7][48,14,14][64,7,7][64,28,28][64,48,72][64,96,144][64,112,112][64,192,288][64,384,576][96,7,7][128,16,20][128,16,40][128,16,80][128,16,160][128,16,2000][128,56,56][128,96,144][16,64,80][16,64,160][16,64,320][16,64,640][16,64,8000][256,8,20][256,8,40][256,8,80][256,8,160][256,8,1000][256,48,72][512,4,20][512,4,40][512,4,80][512,4,160][512,4,500][512,24,36][500,1,32]
kSIGMOID:[1,48,72][1,96,144][1,128,224][1,192,288][1,256,448][1,300,8][1,300,256][1,300,2048][1,3600,256][8,48,72][8,96,144][8,192,288][12,77,2][12,256,2][12,512,2][16,225,768][128,1,1][128,60,60][256,60,60]
kTANH:[1,1,768][10,64,128][128,1,1][3,512,512]
kLEAKY_RELU(alpha=0.3):[12,256,256][32,256,256][64,128,128]
Concatenation
Tensor data types: FP16 or FP32.
The displayed input shapes omit the fixed leading batch dimension. Axis values are full NCHW tensor axes, so
axis=0is the batch axis. The validatedaxis=0cases use one input and therefore preserve batch size1.Up to 40 inputs have been validated.
Validated configurations (input
[C,H,W]per tensor ×num_inputs, axis):axis=0:1×
[1,1,1024]1×
[1,1,32128]1×
[1,64,1024]
axis=1:2×
[128,28,28]2×
[128,7,7]+[32,7,7]2×
[160,7,7]+[64,7,7]2×
[128,14,14]+[384,14,14]2×
[32,7,7]+[64,7,7]2×
[32,7,7]+[96,7,7]3×
[192,60,60]5×
[128,60,60]
axis=2:2×
[128,1,160]2×
[128,1,20]2×
[128,1,40]2×
[128,1,80]
axis=3:2×
[1,1,32]2×
[1,64,32]2×
[1,300,2]2×
[1,3600,2]2×
[12,256,1]2×
[12,512,1]2×
[12,77,1]2×
[64,1,32]2×
[300,64,1]4×
[1,300,128]20×
[128,1,1]40×
[128,1,1]
Constant
Produces a constant tensor consumed by other DLA layers.
Supported value types: FP32, INT32, INT8, UINT8.
FP32: general-purpose values, including scale tensors, bias tensors, and convolution weights.
INT8: valid only as quantized convolution weight kernels consumed by
IConvolutionLayer.UINT8: valid only as the data input to
IDequantizeLayer, or as the zero-point input toIQuantizeLayer/IDequantizeLayer.INT32: valid only as index or shape tensors consumed by layers that accept integer types (for example,
ISliceLayer,IGatherLayer).
Validated
[C,H,W]shapes, grouped by value type:FP32 (4):
[1,1,128][12,64,512][3,1,1][8,300,32]
INT32 (2):
[1,1,1][12,64,64]
INT8 (27):
[1,3,3][3,3,3][4,3,3][8,7,7][16,3,3][24,1,1][24,3,3][32,1,1][32,3,3][48,3,3][60,3,3][64,1,1][64,3,3][64,4,4][64,7,7][96,3,3][128,1,1][128,3,3][128,4,4][132,3,3][160,3,3][224,3,3][256,1,1][256,3,3][512,1,1][512,3,3][1024,1,1]
UINT8 (20):
[1,3,3][3,7,7][3,16,16][64,1,1][64,3,3][96,1,1][128,1,1][128,3,3][192,1,1][256,1,1][256,3,3][512,1,1][512,3,3][576,1,1][640,1,1][768,1,1][1024,1,1][2048,1,1][3072,1,1][4096,1,1]
Convolution
Tensor data types: FP16 or FP32 for the input feature map; FP32 for weights and bias.
2D convolutions only.
Regular convolutions (
num_groups=1) and depthwise convolutions (num_groupsequals the number of input channels) are supported.Dilation must be
[1,1](no dilated convolutions).Stride values up to
[16,16]have been validated (for patch-embedding kernels).Bias mode:
CHANNELonly.Validated configurations, grouped by input
[C,H,W](kernel[C/g,R,S], stride;*denotes depthwise):[3,60,8000]:[3,3,3]s=[1,1][3,224,224]:[3,3,3]s=[1,1];[3,3,3]s=[2,2][3,768,1152]:[3,7,7]s=[2,2][3,960,960]:[3,16,16]s=[16,16][4,128,224]:[4,3,3]s=[2,2][16,16,2000]:[16,3,3]s=[1,1][16,32,4000]:[16,3,3]s=[1,1][24,64,112]:[24,3,3]s=[1,1];[24,3,3]s=[2,2][24,128,224]:[24,3,3]s=[1,1][32,8,1000]:[32,3,3]s=[1,1][32,14,14]:[32,3,3]s=[1,1][32,16,2000]:[32,1,1]s=[1,1][32,112,112]:*[1,3,3]s=[1,1];[32,1,1]s=[1,1];[32,3,3]s=[1,1];[32,3,3]s=[2,2][32,224,224]:[32,3,3]s=[2,2][48,7,7]:[48,3,3]s=[1,1][48,14,14]:[48,3,3]s=[1,1];[48,3,3]s=[2,2][60,32,56]:[60,3,3]s=[1,1];[60,3,3]s=[2,2][64,4,500]:[64,3,3]s=[1,1][64,7,7]:[64,3,3]s=[1,1][64,28,28]:[64,3,3]s=[1,1];[64,3,3]s=[2,2][64,48,72]:[64,1,1]s=[1,1];[64,3,3]s=[1,1][64,56,56]:[64,1,1]s=[1,1];[64,3,3]s=[1,1];[64,3,3]s=[2,2][64,96,144]:[64,1,1]s=[1,1];[64,3,3]s=[1,1][64,112,112]:*[1,3,3]s=[2,2];[64,1,1]s=[1,1][64,192,288]:[64,1,1]s=[1,1];[64,1,1]s=[2,2];[64,3,3]s=[1,1];[64,3,3]s=[2,2][96,7,7]:[96,3,3]s=[1,1][96,16,28]:[96,3,3]s=[1,1];[96,3,3]s=[2,2][128,7,7]:[128,3,3]s=[1,1][128,8,1000]:[128,1,1]s=[1,1][128,16,2000]:[128,1,1]s=[1,1][128,28,28]:[128,1,1]s=[1,1];[128,3,3]s=[1,1];[128,3,3]s=[2,2][128,56,56]:*[1,3,3]s=[1,1];*[1,3,3]s=[2,2];[128,1,1]s=[1,1];[128,3,3]s=[2,2][128,60,60]:[128,3,3]s=[1,1][128,96,144]:[128,1,1]s=[1,1];[128,1,1]s=[2,2];[128,3,3]s=[1,1];[128,3,3]s=[2,2][132,8,14]:[132,3,3]s=[1,1][160,7,7]:[160,3,3]s=[1,1][224,7,7]:[224,3,3]s=[1,1][256,4,500]:[256,1,1]s=[1,1][256,7,7]:[256,3,3]s=[1,1][256,8,1000]:[256,1,1]s=[1,1][256,14,14]:[256,3,3]s=[1,1];[256,3,3]s=[2,2][256,28,28]:*[1,3,3]s=[1,1];[256,1,1]s=[1,1];[256,3,3]s=[1,1];[256,3,3]s=[2,2][256,48,72]:[256,1,1]s=[1,1];[256,1,1]s=[2,2];[256,3,3]s=[1,1];[256,3,3]s=[2,2][512,4,500]:[512,1,1]s=[1,1][512,7,7]:[512,1,1]s=[1,1];[512,3,3]s=[1,1][512,14,14]:*[1,3,3]s=[1,1];[512,1,1]s=[1,1];[512,3,3]s=[2,2][512,24,36]:[512,1,1]s=[1,1];[512,3,3]s=[1,1][576,60,60]:[576,1,1]s=[1,1][640,60,60]:[640,1,1]s=[1,1][768,1,1]:[768,1,1]s=[1,1][1024,1,1]:[1024,1,1]s=[1,1][1024,7,7]:*[1,3,3]s=[1,1][1024,64,1]:[1024,1,1]s=[1,1]
Dequantize
Input: UINT8. Output: FP32 (or FP16 at I/O boundaries).
Axis=1; asymmetric quantization only (zero-point ≠ 0 permitted).
Scale: FP32 scalar per channel. Zero-point: UINT8 scalar per channel.
Validated input
[C,H,W]shapes:[1,1,112][1,1,500][1,1,512][1,1,1000][1,1,2000][1,1,4000][1,1,8000][1,16,4][1,48,72][1,96,144][1,128,224][1,192,288][1,256,448][3,224,224][3,60,8000][3,768,1152][4,128,224][8,48,72][8,96,144][8,192,288][16,4,500][16,16,2000][16,32,4000][16,64,8000][24,32,56][24,64,112][24,128,224][32,7,7][32,8,1000][32,14,14][32,16,2000][32,32,4000][32,112,112][32,224,224][48,7,7][48,14,14][60,16,28][60,32,56][64,7,7][64,4,500][64,24,36][64,28,28][64,48,72][64,56,56][64,96,144][64,112,112][64,192,288][64,384,576][70,28,28][96,7,7][96,8,14][96,16,28][128,8,1000][128,14,14][128,16,2000][128,28,28][128,56,56][128,96,144][132,8,14][256,4,500][256,7,7][256,8,1000][256,14,14][256,28,28][256,48,72][384,14,14][500,1,2][500,1,10][500,1,32][500,1,64][512,1,1][512,4,500][512,7,7][512,14,14][512,24,36]
ElementWise
Tensor data types: FP16 or FP32.
Supported operations (
ElementWiseOperation):kSUM,kSUB,kPROD,kDIV.Comparison operations are supported only in these complete subgraph patterns:
kEQUAL,kGREATER, orkLESSimmediately followed by a cast to FP16 or INT8.kGREATERorkLESSfollowed bykNOTand then a cast to FP16 or INT8.
The INT8 cast in these patterns is a supported intermediate conversion; it does not add general pure-INT8 compute support.
Broadcasting is supported along dimensions whose size is 1.
Validated configurations (input1
[C,H,W]OP input2[C,H,W]):kDIV:[1,1,1024]OP[1,1,1][1,1,768]OP[1,1,1][1,1,768]OP[1,1,768][1,64,1024]OP[1,64,1][256,60,60]OP[1,60,60]
kSUB:[1,300,1]OP[1,300,1][256,60,60]OP[1,60,60]
kPROD:[1,1,100]OP[1,1,100][1,1,1000]OP[256,8,1000][1,1,1000]OP[32,8,1000][1,1,1024]OP[1,1,1024][1,1,2000]OP[128,16,2000][1,1,2000]OP[16,16,2000][1,1,4000]OP[32,32,4000][1,1,500]OP[16,4,500][1,1,500]OP[512,4,500][1,1,500]OP[64,4,500][1,1,768]OP[1,1,768][1,1,8000]OP[16,64,8000][1,300,2048]OP[1,300,2048][1,300,256]OP[1,300,256][1,3600,256]OP[1,3600,256][1,64,1024]OP[1,64,1024][1,64,768]OP[1,64,1][12,256,1]OP[12,256,1][12,512,1]OP[12,512,1][12,77,1]OP[12,77,1][128,60,60]OP[128,60,60][16,225,768]OP[16,225,768][16,300,4]OP[1,300,4][256,60,60]OP[256,60,60][300,64,2]OP[300,1,2]
kSUM:[1,1,64]OP[1,1,64][1,1,768]OP[1,1,768][1,1,1024]OP[1,1,1024][1,49,1024]OP[1,49,1024][1,64,64]OP[1,1,64][1,64,64]OP[1,64,64][1,64,768]OP[1,64,768][1,64,1024]OP[1,64,1024][1,196,512]OP[1,196,512][1,256,768]OP[1,256,768][1,300,1]OP[1,300,1][1,300,2]OP[1,300,2][1,512,768]OP[1,512,768][1,784,256]OP[1,784,256][1,3136,128]OP[1,3136,128][3,512,512]OP[3,512,512][12,64,64]OP[12,64,64][12,77,77]OP[1,1,77][12,77,77]OP[12,77,77][12,256,256]OP[1,1,256][12,256,256]OP[12,256,256][12,512,512]OP[1,1,512][12,512,512]OP[12,512,512][16,225,192]OP[16,225,192][24,64,112]OP[24,64,112][60,32,56]OP[60,32,56][64,1,64]OP[64,1,64][64,48,72]OP[64,48,72][64,96,144]OP[64,96,144][64,112,112]OP[64,112,112][64,128,128]OP[64,128,128][64,192,288]OP[64,192,288][96,16,28]OP[96,16,28][128,56,56]OP[128,56,56][128,96,144]OP[128,96,144][256,16,32]OP[256,16,32][256,28,28]OP[256,28,28][256,48,72]OP[256,48,72][300,1,2]OP[300,64,2][32,256,256]OP[32,256,256][512,7,7]OP[512,7,7][512,14,14]OP[512,14,14][512,24,36]OP[512,24,36]
Normalization
Tensor data types: FP16.
Layer normalization only. The axes bitmask must cover only the last dimension (bit 3) or the last two dimensions (bits 2 and 3).
Scale and bias tensors must be build-time constants whose size matches the normalized dimension(s).
Epsilon values in the range
[1e-7, 1e-5]have been validated.Validated
[C,H,W]shapes:axes=[3](last dim, 12 shapes):[1,1,1024][1,49,1024][1,64,768][1,77,768][1,196,512][1,256,768][1,300,256][1,512,768][1,784,256][1,3136,128][1,3600,256][16,225,192]
axes=[2,3](last two dims, 3 shapes):[64,64,128][128,32,64][256,16,32]
Pooling
Tensor data types: FP16 or FP32.
Supported types (
PoolingType):kMAX,kAVERAGE.Maximum window size: 7×7.
Validated configurations (input
[C,H,W], window, stride):kAVERAGE:[512,7,7]win=[7,7]stride=[2,2]
kMAX:[16,64,80]win=[2,2]stride=[2,2][16,64,160]win=[2,2]stride=[2,2][16,64,320]win=[2,2]stride=[2,2][16,64,640]win=[2,2]stride=[2,2][16,64,8000]win=[2,2]stride=[2,2][32,32,40]win=[2,2]stride=[2,2][32,32,80]win=[2,2]stride=[2,2][32,32,160]win=[2,2]stride=[2,2][32,32,320]win=[2,2]stride=[2,2][32,32,4000]win=[2,2]stride=[2,2][64,384,576]win=[3,3]stride=[2,2][128,16,20]win=[2,1]stride=[2,1][128,16,40]win=[2,1]stride=[2,1][128,16,80]win=[2,1]stride=[2,1][128,16,160]win=[2,1]stride=[2,1][128,16,2000]win=[2,2]stride=[2,2][256,8,20]win=[2,1]stride=[2,1][256,8,40]win=[2,1]stride=[2,1][256,8,80]win=[2,1]stride=[2,1][256,8,160]win=[2,1]stride=[2,1][256,8,1000]win=[2,2]stride=[2,2]
Quantize
Input: FP16. Output: UINT8.
Axis=1; asymmetric quantization only (zero-point ≠ 0 permitted).
Scale: FP32 scalar per channel. Zero-point: UINT8 scalar per channel.
Validated input
[C,H,W]shapes:[1,1,112][1,1,500][1,1,512][1,1,1000][1,1,2000][1,1,4000][1,4,28][1,16,4][1,48,72][1,96,144][1,128,224][1,192,288][1,256,448][3,224,224][3,60,8000][3,768,1152][4,128,224][8,48,72][8,96,144][8,192,288][16,4,500][16,16,2000][16,32,4000][16,64,8000][24,32,56][24,64,112][24,128,224][28,28,70][32,7,7][32,8,1000][32,14,14][32,16,2000][32,32,4000][32,112,112][32,224,224][48,7,7][48,14,14][60,16,28][60,32,56][64,7,7][64,4,500][64,24,36][64,28,28][64,48,72][64,56,56][64,96,144][64,112,112][64,192,288][64,384,576][70,28,28][96,7,7][96,8,14][96,16,28][128,8,1000][128,14,14][128,16,2000][128,28,28][128,56,56][128,96,144][128,224,1][132,8,14][256,4,500][256,7,7][256,8,1000][256,14,14][256,28,28][256,48,72][256,448,1][384,14,14][500,1,2][500,1,10][500,1,32][500,1,64][512,1,1][512,4,500][512,7,7][512,14,14][512,24,36]
Shuffle
Tensor data types: FP16 or FP32.
IShuffleLayercovers reshape, transpose, depth-to-space, and space-to-depth operations. The DLA Compiler test source identifies all four patterns, but this page has no exact validated shape and attribute tuples for them. Treat these patterns as unconfirmed until the target engine builds successfully.
I/O Formats on DLA#
TensorRT supports the following tensor formats for network inputs and outputs.
Specify formats with the TensorFormat enum on the network’s input and
output bindings.
Data type |
|
Notes |
|---|---|---|
FP16 |
|
Supported for both inputs and outputs. |
FP32 |
|
Supported for both inputs and outputs. |
UINT8 |
|
Supported for outputs only. Requires the network to end with an
|
FP16/kCHW16 and FP32/kLINEAR are validated as inputs and outputs in the
test suite. UINT8/kLINEAR is validated as an output only.
Note
TensorFormat::kDLA_LINEAR is a DLA-specific planar layout distinct from
kLINEAR. Specify kLINEAR at the TensorRT API level. TensorRT maps it
to the correct layout internally.
For comparison with Jetson and DriveOS DLA formats, refer to I/O Formats on DLA.
Strong Typing#
Important
All networks are strongly typed, including DLA networks. TensorRT 11.0 removed the weak-typing APIs, so DLA on Windows on ARM requires strongly typed networks. The following builder flags are not valid for DLA builds:
BuilderFlag::kFP16BuilderFlag::kINT8BuilderFlag::kPREFER_PRECISION_CONSTRAINTSBuilderFlag::kOBEY_PRECISION_CONSTRAINTS
Use tensor types in the network (or a pre-quantized ONNX import) to set precision. Refer to Strongly Typed Networks for DLA, ONNX Parser Flags for DLA, and Migrating from TensorRT 10.x to 11.x.