Deviceless Ahead-of-time Compilation
Deviceless Ahead-of-time Compilation
Ahead-of-time Compilation
Since cuDNN 9.8, customers are allowed to create and finalize an execution plan devicelessly with a device property descriptor. This helps customers to cover the plan build time ahead of the execution.
Typical workflow:
- Create a device property descriptor from the device, serialize it out. This requires the device.
- Deserialize the device property, create an execution plan from it as well as the computation graph, serialize the plan out. This doesn’t require the device.
- Deserialize and execute the execution plan on devices with the same properties. This requires the device.
Refer to the corresponding C++ sample in samples/cpp/misc/deviceless_aot_compilation.cpp.
What “deviceless” means at each phase
One shared, read-only DeviceProperties descriptor can be passed concurrently to deserialize from multiple threads, enabling efficient parallel deserialization of cached plans without creating per-thread cuDNN handles. (Each deserialize produces an independent Graph object; they do not share mutable state.)
C++ API
The deserialize(blob) overload (no handle) is available since cuDNN 9.8 at the API level.
Runtime test coverage follows the existing deviceless sample policy and gates at 9.11.
Python API
cuDNN Device Properties
cuDNN device property descriptor describes the properties of a GPU device, is serializable and can be used to query cuDNN heuristics / create an execution plan directly without the device to be available.
The API to create a device property descriptor is:
Ways to initialize the device properties:
The API to set a device property descriptor is: