Other Methods#

Frontend to Backend Traits Conversion#

cufftdx::experimental::utils::frontend_to_backend(...) converts a frontend FFT description to a effective backend traits that is used to query the cuFFTDx databases.

In combination with cuFFT Device API, this function can be used to generate the LTO database containing both device function code and metadata (a C++ header file) for a specified FFT operation. See Custom LTO Helper for an example.

#include "cufftdx/utils.hpp"

namespace cufftdx {
   namespace utils {
      enum class algorithm {
         ct,
         bluestein,
      };

      enum class execution_type {
         thread,
         block,
      };
   } // namespace utils

   namespace experimental {
      namespace utils {
         struct backend_impl_traits {
            unsigned int  size;
            fft_type      type;
            fft_direction direction;
            unsigned int  sm;
            unsigned int  elements_per_thread;
            unsigned int  min_elements_per_thread;
         };

         backend_impl_traits backend_traits =
            frontend_to_backend(
               cufftdx::utils::algorithm algo,
               cufftdx::utils::execution_type exec_type,
               unsigned int fft_size
                   /* size_of<FFT>::value */,
               fft_type type
                   /* type_of<FFT>::value */,
               fft_direction direction
                   /* direction_of<FFT>::value */,
               unsigned int sm
                   /* sm_of<FFT>::value */,
               real_mode real_mode
                   /* real_mode_of<FFT>::value */,
               unsigned int elements_per_thread
                   /* elements_per_thread_of<FFT>::value or
                    * 0 if not set */,
               unsigned int block_dim_x
                   /* block_dim_of<FFT>::x or 0 if not set */,
               experimental::code_type code_type
                   /* experimental::code_type_of<FFT>::value */
            );
      } // namespace utils
   } // namespace experimental
} // namespace cufftdx

Runtime Database Query#

cufftdx::experimental::utils::query_database(...) and cufftdx::experimental::utils::get_all_implementations(...) provide runtime access to the cuFFTDx internal database. Both methods are useful when FFT configuration is selected dynamically, for example based on user inputs. It can be used to check if a given FFT configuration is supported without workspace. To better understand the difference between non workspace and workspace required FFTs please check Supported Functionality. cufftdx::experimental::utils::query_database(...) returns the optimal implementation as std::optional<cufftdx::experimental::utils::frontend_impl_traits>. If no implementation is found for the provided configuration, it returns std::nullopt.

The code_type parameter is of type cufftdx::experimental::query_code_type and selects which database backend the call dispatches to:

  • experimental::query_code_type::ptx – query the built-in offline PTX database that ships with cuFFTDx (default).

  • experimental::query_code_type::ltoir_offline – query an offline LTO database that has been generated ahead of time with the LTO helper (see Custom LTO Helper) and included in the translation unit.

  • experimental::query_code_type::ltoir_online – forward the request to the cuFFT Device API at runtime via query_all_cufft_implementations. This mode additionally requires CUFFTDX_ENABLE_CUFFT_DEPENDENCY to be defined and the application to be linked against cuFFT.

std::optional<cufftdx::experimental::utils::frontend_impl_traits>
cufftdx::experimental::utils::query_database(
    unsigned int                    fft_size,
    cufftdx::fft_direction          dir,
    cufftdx::fft_type               type,
    unsigned int                    sm,
    cufftdx::utils::execution_type  execution,
    cufftdx::precision              prec =
        cufftdx::precision::f32,
    unsigned int                    fft_ept = 0,
    unsigned int                    ffts_per_block = 0,
    std::tuple<unsigned int, unsigned int, unsigned int>
        block_dim = {0, 0, 0},
    cufftdx::complex_layout         layout =
        cufftdx::complex_layout::natural,
    cufftdx::real_mode              rmode =
        cufftdx::real_mode::normal,
    cufftdx::experimental::query_code_type code_type =
        cufftdx::experimental::query_code_type::ptx);

cufftdx::experimental::utils::get_all_implementations(...) returns all matching implementations as std::vector<cufftdx::experimental::utils::frontend_impl_traits>. If no implementation is found, it returns an empty vector.

std::vector<cufftdx::experimental::utils::frontend_impl_traits>
cufftdx::experimental::utils::get_all_implementations(
    unsigned int                    fft_size,
    cufftdx::fft_direction          dir,
    cufftdx::fft_type               type,
    unsigned int                    sm,
    cufftdx::utils::execution_type  execution,
    cufftdx::precision              prec =
        cufftdx::precision::f32,
    unsigned int                    fft_ept = 0,
    unsigned int                    ffts_per_block = 0,
    std::tuple<unsigned int, unsigned int, unsigned int>
        block_dim = {0, 0, 0},
    cufftdx::complex_layout         layout =
        cufftdx::complex_layout::natural,
    cufftdx::real_mode              rmode =
        cufftdx::real_mode::normal,
    cufftdx::experimental::query_code_type code_type =
        cufftdx::experimental::query_code_type::ptx);

For an example on how to use these APIs, please check the nvrtc_query_database_fft_block example.

Note

These APIs are available only when CUFFTDX_ENABLE_RUNTIME_DATABASE is defined. The default experimental::query_code_type::ptx mode only sees FFT sizes present in the PTX database that ships with cuFFTDx (i.e. those that do not require a workspace). To extend the queryable set of sizes, use one of the LTO modes: experimental::query_code_type::ltoir_offline (after generating an offline LTO database with the LTO helper, see Custom LTO Helper) or experimental::query_code_type::ltoir_online (which delegates to the cuFFT Device API at runtime and additionally requires CUFFTDX_ENABLE_CUFFT_DEPENDENCY).

The LTO query modes are only supported for LTO databases generated with cuFFT 12.3 or later (from CUDA Toolkit 13.3).

Warning

Compilation time can significantly increase when using these APIs.

(online) LTO Database Creation#

cufftdx::utils::get_database_and_ltoir() returns a tuple containing:

  • Database string (std::string) to be inserted into the cuFFTDx headers.

  • Vector of LTOIRs (std::vector<std::vector<char>>) for building device functions for the specified FFT operation.

  • Required CUDA block dimensions (dim3) for executing the FFT operation.

  • Required shared memory size (unsigned, in bytes) for executing the FFT operation.

std::tuple<std::string, std::vector<std::vector<char>>, dim3,
           unsigned int>
cufftdx::utils::get_database_and_ltoir(
    unsigned int                    fft_size,
    cufftdx::fft_direction          dir,
    cufftdx::fft_type               type,
    unsigned int                    sm,
    cufftdx::utils::execution_type  execution,
    cufftdx::precision              prec =
        cufftdx::precision::f32,
    cufftdx::complex_layout         layout =
        cufftdx::complex_layout::natural,
    cufftdx::real_mode              rmode =
        cufftdx::real_mode::normal,
    unsigned int                    fft_ept = 0
        /* use heuristic */,
    unsigned int                    ffts_per_block = 1
        /* 0: use suggested ffts_per_block */,
    dim3                            block_dim = dim3(0, 0, 0)
        /* 0: use suggested block_dim */);

Example

// Assuming the following FFT operator is defined in the NVRTC-compiled code:
// using FFT = decltype(cufftdx::Block() +
//                      cufftdx::Size<128>() +
//                      cufftdx::Type<cufftdx::fft_type::c2c>() +
//                      cufftdx::Direction<cufftdx::fft_direction::forward>() +
//                      cufftdx::Precision<float>() +
//                      cufftdx::ElementsPerThread<8>() +
//                      cufftdx::FFTsPerBlock<2>() +
//                      cufftdx::SM<750>());

// You can get the database string and LTOIRs for the FFT operation by calling:
auto [lto_db, ltoirs, block_dim, sm_size] =
   cufftdx::utils::get_database_and_ltoir(128,
                                          cufftdx::fft_direction::forward,
                                          cufftdx::fft_type::c2c,
                                          750,
                                          cufftdx::utils::execution_type::block,
                                          cufftdx::precision::f32,
                                          cufftdx::complex_layout::natural,
                                          cufftdx::real_mode::normal,
                                          8,
                                          2);

After obtaining the database string and LTOIRs, you can insert the database string into the cuFFTDx header file and link the LTOIRs to the user code as shown in Use Case II: Online Kernel Generation.

Note

The cufftdx::utils::get_database_and_ltoir() function is a wrapper around cuFFT Device APIs (see cuFFT Device API Reference). To use this function:

  1. Define CUFFTDX_ENABLE_CUFFT_DEPENDENCY.

  2. Link against the cuFFT library.

Querying All cuFFT Online LTO Implementations#

cufftdx::utils::query_all_cufft_implementations() is the “list all candidates” counterpart to get_database_and_ltoir. It queries the cuFFT Device API for every LTOIR implementation that matches the requested FFT configuration and returns them as a std::vector<cufftdx::experimental::utils::frontend_impl_traits>. Each returned entry has a distinct elements-per-thread value, so the caller can inspect the full set of valid choices (for example, to autotune across them) before calling get_database_and_ltoir() for the chosen configuration.

std::vector<cufftdx::experimental::utils::frontend_impl_traits>
cufftdx::utils::query_all_cufft_implementations(
    unsigned int                    fft_size,
    cufftdx::fft_direction          dir,
    cufftdx::fft_type               type,
    unsigned int                    sm,
    cufftdx::utils::execution_type  execution,
    cufftdx::precision              prec =
        cufftdx::precision::f32,
    cufftdx::complex_layout         layout =
        cufftdx::complex_layout::natural,
    cufftdx::real_mode              rmode =
        cufftdx::real_mode::normal,
    unsigned int                    fft_ept = 0
        /* use heuristic */,
    unsigned int                    ffts_per_block = 1
        /* 0: use suggested ffts_per_block */,
    std::tuple<unsigned int, unsigned int, unsigned int>
        block_dim = {0, 0, 0});

This function is the backend for query_code_type::ltoir_online in query_database / get_all_implementations. For a complete autotuning example that enumerates all implementations, benchmarks each one, and then calls get_database_and_ltoir() for the selected configuration, see example/cufftdx/04_nvrtc_fft/nvrtc_query_autotune_lto.cu.

Note

Like get_database_and_ltoir(), this function is a wrapper around the cuFFT Device API and requires CUFFTDX_ENABLE_CUFFT_DEPENDENCY to be defined and the application to be linked against cuFFT.

Shared Memory Compute For Dynamic Batching#

cufftdx::experimental::utils::get_shared_memory_size_for_dynamic_batching<FFT>(const unsigned int ffts_per_block) computes the total shared memory bytes required for computing multiple FFTs with the DynamicBatching operator.

template<class FFT>
constexpr unsigned int
get_shared_memory_size_for_dynamic_batching(
    const unsigned int ffts_per_block);

This function is useful when you need to determine shared memory requirements at runtime, particularly when the number of FFTs per block varies dynamically.

Parameters:

  • ffts_per_block - Number of FFTs to compute per block (user-defined, can be a runtime value)

cufftdx::experimental::utils::get_shared_memory_size_for_dynamic_batching(const unsigned int shared_memory_size_per_fft, const unsigned int ffts_per_block, const unsigned int implicit_type_batching) provides the same functionality at runtime, without creating a FFT Description.

constexpr unsigned int
get_shared_memory_size_for_dynamic_batching(
    const unsigned int shared_memory_size_per_fft,
    const unsigned int ffts_per_block,
    const unsigned int implicit_type_batching);

Parameters:

  • shared_memory_size_per_fft - Shared memory size required for a single FFT (obtained from FFT::shared_memory_size)

  • ffts_per_block - Number of FFTs to compute per block (user-defined, can be a runtime value)

  • implicit_type_batching - Number of FFTs batched per value type (obtained from FFT::implicit_type_batching)

Example

using FFT = decltype(Size<128>() + Precision<float>() +
               Type<fft_type::c2c>() +
               Direction<fft_direction::forward>() +
               ElementsPerThread<8>() + DynamicBatching());

// Query compile-time traits
constexpr auto shared_memory_size_per_fft =
    FFT::shared_memory_size;
constexpr auto implicit_type_batching =
    FFT::implicit_type_batching;

// Runtime value chosen by user
unsigned int ffts_per_block = 8;

// Compute total shared memory requirement
auto total_shared_memory_size_bytes =
    cufftdx::experimental::utils::
        get_shared_memory_size_for_dynamic_batching(
            shared_memory_size_per_fft,
            ffts_per_block,
            implicit_type_batching);

auto total_shared_memory_size_bytes_with_desc =
    cufftdx::experimental::utils::
        get_shared_memory_size_for_dynamic_batching<FFT>(
            ffts_per_block);

Note

These functions only return a valid result when using traits from a description with the DynamicBatching operator. For obtaining the shared memory bytes required for other descriptions, see Shared Memory Size Trait.