Blacklisting and ECC Error Recovery#
Pages that have been previously retired are blacklisted for all future allocations of the framebuffer, provided that the target GPU has been properly reattached and initialized. This chapter presents a procedure for ensuring that retired pages are blacklisted and all GPUs have recovered from the ECC error.
Note
This procedure requires the termination of all clients on the target GPU. It is not possible to blacklist a new page while clients remain active.
Blacklisting: Procedure Overview:
Verify that there are pending retired pages.
Determine GPUs that are related and must be reattached together.
Stop applications and verify there are none left running.
Reinitialize the GPU, or reboot the system.
Verify that blacklisting has occurred.
Restart applications.
Verifying Retired Pages are Pending#
When pages are retired but have not yet been blacklisted, the retired pages are marked as pending for that GPU. This can be seen through nvidia-smi:
$ nvidia-smi -i <target gpu> -q -d PAGE_RETIREMENT
...
Retired pages
Single Bit ECC : 2
Double Bit ECC : 0
Pending Page Blacklist : Yes
...
If Pending Page Blacklist shows “No”, then all retired pages have already been blacklisted.
If Pending Page Blacklist shows “Yes”, then at least one of the retired pages that are counted are not yet blacklisted. Note that the exact count of pending pages is not shown.
Note
The retired pages count increments immediately when a page is retired and not on the next driver reload when the page is blacklisted.
Stopping GPU Clients#
Before the the NVIDIA driver can reattach the GPUs, clients that are using those GPUs must first be stopped and the GPU must be unused.
All applications that are using the GPUs should first be stopped. Use nvidia-smi to list processes that are actively using the GPUs. In
the example below, a tensorflow python program is using both GPUs 0 and 1. Both will need to be stopped.
$ nvidia-smi
...
+---------------------------------------------------------------------+
| Processes: GPU Memory |
| GPU PID Type Process name Usage |
|=====================================================================|
| 0 8962 C python 15465MiB |
| 1 8963 C python 15467MiB |
+---------------------------------------------------------------------+
Once all applications are stopped, nvidia-smi should show no processes found:
$ nvidia-smi
...
+---------------------------------------------------------------------+
| Processes: GPU Memory |
| GPU PID Type Process name Usage |
|=====================================================================|
| No running processes found |
+---------------------------------------------------------------------+
On Linux systems, additional software infrastructure can hold the GPU open and prevent the GPU from being detached by the driver. These
include the nvidia-persistenced, and version 1 of nvidia-docker. Nvidia-docker version 2 does not need to be stopped.
A list of open proesses using the driver can be verified on Linux with the lsof command:
$ sudo lsof /dev/nvidia*
COMMAND PID USER FD TYPE DEVICE NODE NAME
nvidia-pe 941 nvidia-persistenced 2u CHR 195,255 453 /dev/nvidiactl
nvidia-pe 941 nvidia-persistenced 3u CHR 195,0 454 /dev/nvidia0
nvidia-pe 941 nvidia-persistenced 4u CHR 195,254 607 /dev/nvidia-modeset
nvidia-pe 941 nvidia-persistenced 5u CHR 195,1 584 /dev/nvidia1
nvidia-pe 941 nvidia-persistenced 6u CHR 195,254 607 /dev/nvidia-modeset
Once all clients of the GPU are stopped, lsof should return no entries:
$ sudo service nvidia-persistenced stop
$ sudo lsof /dev/nvidia*
$
Reattaching the GPU#
Reattaching the GPU, to blacklist pending retired pages, can be done in several ways. In order of cost, from low to high:
Re-attach the GPUs (persistence mode disabled only)
Reset the GPUs
Reload the kernel module (nvidia.ko)
Reboot the machine (or VM)
Reattaching the GPU is the least invasive solution. The detachment process occurs automatically a few seconds after the last client terminates on the GPU, as long as persistence mode is not enabled. The next client that targets the GPU will trigger the driver to reattach and blacklist all marked pages.
If persistence mode is enabled, the preferred solution is to reset the GPU using nvidia-smi. To reset an individual GPU:
$ nvidia-smi -i < target GPU> -r
Or to reset all GPUs together:
$ nvidia-smi -r
These operations reattach the GPU as a step in the larger process of resetting all GPU SW and HW state.
Reloading the NVIDIA kernel module triggers reattachment of all GPUs on the machine, and thus requires the termination of all clients on all GPUs.
Finally, rebooting the machine will effectively reattach the GPUs as the driver is reloaded and reinitialized during reboot. While rebooting isn’t required and is highly invasive, it might simplify the recovery action in some operating environments.