GDS
Overview
GDS (GPUDirect Storage) enables direct transfers between GPU memory and files without CPU bounce buffers, using the NVIDIA cuFile API. It uses cuFile batch transfers by default and can select the multi-threaded engine with mode=mt. The default behavior of createBackend("GDS") remains unchanged. This backend requires GDS-capable hardware and drivers.
Installation
Prerequisites
- A GPU with GPUDirect Storage support (NVIDIA data center GPUs)
- CUDA Toolkit 11.4 or later installed
- cuFile driver and libraries (included with CUDA Toolkit 11.4+)
- A compatible filesystem (ext4, XFS, or a parallel filesystem with GDS support)
Verify GDS Installation
The cuFile libraries required for GDS are included with the CUDA Toolkit. Verify they are present:
You should see libcufile.so and related library files.
Run the GDS compatibility check:
This verifies GPU support, kernel driver compatibility, and filesystem readiness.
Configuration
Backend Parameters
Environment Variables
Build Options
When to Use
- Direct GPU-to-file on NVMe storage — Bypass CPU bounce buffers for maximum throughput on NVMe drives.
- Checkpoint save/load from GPU memory — Write GPU tensors directly to storage without staging through host memory.
- Eliminating CPU bounce buffer overhead — Remove the CPU memory copy step in GPU-to-file transfers.
- Batch I/O by default — Use
GDSwithout parameters, or withmode=batch, for cuFile batch transfers. - Multi-threaded I/O — Use
GDSwithmode=mtto distribute I/O across persistent TaskFlow workers.