GPU Operational Event Catalog#
Warning
GPU Operational Events are a preview feature and remain under active development. Category values, module signatures, module event codes, severities, context payloads, and the set of published events may change before the interface is declared stable. Treat the identifiers in this document as valid for the driver release documented here, and re-validate consumers against each driver release. Legacy Xid output remains the stable contract for production monitoring during the preview period.
This page lists the shared event category registry and the GPU Operational Events published in this driver release. See Analyzing GPU Operational Events for monitoring interfaces and how to interpret an event, and GPU Operational Event Formats for the formats events are published in.
Event Categories#
The category registry is shared by all reporting modules. A consumer can act on the category alone when it does not recognize a specific module event code.
Category |
Value |
Indicates |
|---|---|---|
|
|
An operation did not complete within its expected time budget. |
|
|
A data integrity error in memory or cache, such as ECC, parity, or poison. |
|
|
An error on a link between the GPU and another device, such as NVLink, PCIe, or C2C. |
|
|
Corrupted driver or device state was detected. |
|
|
A security policy violation or an authentication failure. |
|
|
A power delivery or power management fault. |
|
|
A thermal limit was reached or a thermal sensor failed. |
|
|
A fault detected by firmware running on the device. |
|
|
An error reported by a GPU engine. |
|
|
A repair, spare, or tracking resource has been used up. |
|
|
An address translation or MMU fault. |
|
|
An error that has not yet been assigned a specific category. |
|
|
Component initialization progress or completion. |
|
|
Component shutdown or teardown. |
|
|
An operational state transition. |
|
|
A configuration change. |
|
|
A resource has been marked for retirement or repair. |
|
|
A resource usage threshold was crossed. |
|
|
Operational telemetry. |
|
|
Engineering and debug data. |
This release publishes events in eight of these categories: TIMEOUT,
MEMORY_INTEGRITY_ERROR, INTERCONNECT_ERROR, FIRMWARE_FAULT, RESOURCE_EXHAUSTED,
UNCLASSIFIED_ERROR, INITIALIZATION, and RESOURCE_RETIREMENT. The remaining values are
defined so that consumers can decode events introduced in later releases.
Note
In this release, the nvidia-smi event-log decoder renders category 0x0008 as
“Engine Software Error”. The numeric value is correct and corresponds to FIRMWARE_FAULT,
so decode categories from the numeric value.
Event Log Levels#
The log level describes the urgency of the report independently from the event severity. Values increase with urgency, so a subscription filter can use “at least this urgent” semantics. A filter value of 0 matches every log level.
Log Level |
Value |
Indicates |
|---|---|---|
|
|
Operational telemetry. |
|
|
Diagnostic information. NVML names this level
|
|
|
A notable condition that does not require immediate action. |
|
|
A condition that may require operator attention. |
|
|
A condition that requires operator attention. |
Event Severities#
The severity identifies the impact of the event relative to its scope. Values increase with severity, so a subscription filter can use “at least this severe” semantics. A filter value of 0 matches every severity.
Severity |
Value |
Indicates |
|---|---|---|
|
|
The event reports state, telemetry, or progress. |
|
|
The condition was corrected without recovery action by the operator. |
|
|
The condition affected operation, and recovery is possible. |
|
|
The condition caused loss of the scoped resource. |
When a GPU Operational Event is encoded as CPER, the GOE severity is remapped to a host-scoped CPER severity. See GOE Severity to CPER Severity Mapping.
Event Scopes#
The scope identifies the resource affected by the event.
Scope |
Value |
Indicates |
|---|---|---|
|
|
The whole system. |
|
|
One GPU device. |
|
|
One MIG GPU instance. |
|
|
One MIG compute instance. |
|
|
One GPU engine. |
Published Events#
This release publishes the events listed below. To identify an event you have decoded, find its module signature and module event code in the lookup table, then follow the event name to its entry. Not all Xids emitted by the driver have been fully converted to GPU Operational Events yet; more support will be added in future releases. To determine the actions to take, look up the event’s legacy Xid in the Xid Errors catalog.
A severity listed as FATAL or RECOVERABLE is selected at runtime according to the impact
of the reported condition.
Event Lookup#
Module |
Code |
Event |
Legacy Xid |
Category |
|---|---|---|---|---|
|
|
None |
|
|
|
|
179 |
|
|
|
|
143 |
|
|
|
|
143 |
|
|
|
|
119 |
|
|
|
|
120 |
|
|
|
|
120 |
|
|
|
|
120 |
|
|
|
|
140 |
|
|
|
|
175 |
|
|
|
|
136 |
|
|
|
|
150 |
|
|
|
|
150 |
|
|
|
|
155 |
|
|
|
|
63 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
64 |
|
|
|
|
160 |
|
|
|
|
161 |
|
|
|
|
160 |
|
|
|
|
161 |
|
|
|
|
156 |
|
|
|
|
156 |
|
|
|
|
157 |
|
|
|
|
157 |
|
|
|
|
140 |
|
|
|
|
92 |
|
|
|
|
92 |
|
|
|
|
177 |
|
|
|
|
Varies |
|
Event Details#
Entries are grouped by source module signature, then ordered by category value and module event code.
GPU#
DRIVER_INIT#
- Code:
0x0001- Category:
INITIALIZATION- Severity:
INFORMATIONAL- Scope:
DEVICE- Legacy Xid:
None
GPU driver initialization completed for the device. Carries driver and device metadata.
GPU-BUS#
C2C_CONTAINMENT#
- Code:
0x0002- Category:
INTERCONNECT_ERROR- Severity:
FATALorRECOVERABLE- Scope:
DEVICE- Legacy Xid:
179
C2C containment was triggered on the chip-to-chip link to the host CPU. Carries the containment code.
GPU-FSP#
BOOT_TIMEOUT#
- Code:
0x0001- Category:
TIMEOUT- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
143
The driver timed out polling for FSP boot completion during GPU initialization.
FUSE_ERROR#
- Code:
0x0002- Category:
FIRMWARE_FAULT- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
143
The FSP fuse error check failed. Carries the fuse status.
GPU-GSP#
RPC_TIMEOUT#
- Code:
0x0001- Category:
TIMEOUT- Severity:
FATALorRECOVERABLE- Scope:
DEVICE- Legacy Xid:
119
The driver timed out waiting for an RPC response from GSP. Carries the RPC function, name, sequence number, and wait duration.
LIBOS_HEARTBEAT_TIMEOUT#
- Code:
0x0002- Category:
TIMEOUT- Severity:
FATALorRECOVERABLE- Scope:
DEVICE- Legacy Xid:
120
The LibOS heartbeat from the GSP microcontroller stopped.
GSP_RM_HEARTBEAT_TIMEOUT#
- Code:
0x0003- Category:
TIMEOUT- Severity:
FATALorRECOVERABLE- Scope:
DEVICE- Legacy Xid:
120
The GSP-RM heartbeat stopped.
POISON#
- Code:
0x0005- Category:
MEMORY_INTEGRITY_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
140
GSP consumed poisoned data.
FIRMWARE_FAULT#
- Code:
0x0004- Category:
FIRMWARE_FAULT- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
120
GSP firmware reported a fault.
GPU-MEMSYS#
MAINT_OP_ERROR#
- Code:
0x0001- Category:
MEMORY_INTEGRITY_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
175
A GPU memory subsystem maintenance operation timed out. Carries the operation type.
GPU-NVLINK#
ALI_TRAINING_FAILURE#
- Code:
0x0001- Category:
INTERCONNECT_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
136
NVLink ALI link training failed. Carries the affected link IDs and debug registers.
SW_LINK_DOWN#
- Code:
0x0004- Category:
INTERCONNECT_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
155
An NVLink went down while it was active. Carries the link ID.
MSE_DEGRADED#
- Code:
0x0002- Category:
FIRMWARE_FAULT- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
150
The NVLink MSE transport entered a degraded state. The driver stops MSE communication until it is reloaded.
MSE_WATCHDOG_TIMEOUT#
- Code:
0x0003- Category:
FIRMWARE_FAULT- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
150
The NVLink MSE watchdog expired. Carries the error status and debug registers.
GPU-RAS#
ECC_RESIDUAL_UNCORRECTABLE_ERROR#
- Code:
0x0013- Category:
MEMORY_INTEGRITY_ERROR- Severity:
FATALorRECOVERABLE- Scope:
DEVICE- Legacy Xid:
140
An uncorrectable ECC error was detected outside the normal handling path. Carries DRAM, LTC, MMU, and PCIe error counts.
DRAM_ECC_INTR_STORM#
- Code:
0x0014- Category:
MEMORY_INTEGRITY_ERROR- Severity:
CORRECTED- Scope:
DEVICE- Legacy Xid:
92
Single-bit ECC error interrupts were disabled on a framebuffer partition because of a high error rate. Carries the partition index.
SM_ECC_INTR_STORM#
- Code:
0x0015- Category:
MEMORY_INTEGRITY_ERROR- Severity:
CORRECTED- Scope:
DEVICE- Legacy Xid:
92
An SM single-bit ECC interrupt storm was detected.
ROW_REMAPPING_FAILURE_TABLE_FULL#
- Code:
0x0002- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
Row remapping failed because the row remapping table is full.
ROW_REMAPPING_FAILURE_BANK_FULL#
- Code:
0x0003- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
Row remapping failed because all reserved rows for the bank are already remapped.
ROW_REMAPPING_FAILURE_RESERVED_ROW#
- Code:
0x0005- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
Row remapping failed because the target is a reserved row.
DRAM_RETIREMENT_FAILURE_INFOROM_FULL#
- Code:
0x0008- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
DRAM retirement failed because the InfoROM object is full.
DRAM_RETIREMENT_FAILURE_HW_LIMIT#
- Code:
0x0009- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
DRAM retirement failed because a hardware limit was reached.
DRAM_RETIREMENT_FAILURE_NO_SPARE#
- Code:
0x000A- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
DRAM retirement failed because no spare remains for retirement.
LTS_REPAIR_FAILURE#
- Code:
0x000C- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
161
L2 slice repair failed because no spare L2 slices remain.
MEMORY_CHANNEL_REPAIR_FAILURE#
- Code:
0x000E- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
161
Memory channel repair failed because no spare channels remain.
TPC_REPAIR_FAILURE_NO_SPARE#
- Code:
0x0011- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
157
TPC retirement failed because no spare TPCs are available.
TPC_REPAIR_FAILURE_NO_SPARE_MIG#
- Code:
0x0012- Category:
RESOURCE_EXHAUSTED- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
157
TPC retirement failed in MIG mode because no spare TPCs remain in the same GPC.
ROW_REMAPPING_FAILURE_INTERNAL_ERROR#
- Code:
0x0004- Category:
UNCLASSIFIED_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
Row remapping failed due to an internal driver error. Carries an error string.
ROW_REMAPPING_FAILURE_PAGE_OFFLINE_FAILURE#
- Code:
0x0006- Category:
UNCLASSIFIED_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
The row is already pending remapping. The remap occurs on the next GPU reset.
DRAM_RETIREMENT_FAILURE_INTERNAL_ERROR#
- Code:
0x0007- Category:
UNCLASSIFIED_ERROR- Severity:
FATAL- Scope:
DEVICE- Legacy Xid:
64
DRAM retirement failed due to an internal driver error. The context identifies the specific reason.
ROW_REMAPPING_PENDING#
- Code:
0x0001- Category:
RESOURCE_RETIREMENT- Severity:
RECOVERABLE- Scope:
DEVICE- Legacy Xid:
63
A new DRAM row was marked for remapping. The remap activates on the next GPU reset.
LTS_REPAIR_PENDING#
- Code:
0x000B- Category:
RESOURCE_RETIREMENT- Severity:
RECOVERABLE- Scope:
DEVICE- Legacy Xid:
160
An L2 slice and its pair were marked for repair. Carries the location and the recovery action needed to activate the repair.
MEMORY_CHANNEL_REPAIR_PENDING#
- Code:
0x000D- Category:
RESOURCE_RETIREMENT- Severity:
RECOVERABLE- Scope:
DEVICE- Legacy Xid:
160
A memory channel and its pair were marked for repair. Carries the location and the recovery action needed to activate the repair.
TPC_REPAIR_PENDING_SAME_GPC#
- Code:
0x000F- Category:
RESOURCE_RETIREMENT- Severity:
RECOVERABLE- Scope:
DEVICE- Legacy Xid:
156
A TPC was retired and replaced with a spare from the same GPC.
TPC_REPAIR_PENDING_DIFFERENT_GPC#
- Code:
0x0010- Category:
RESOURCE_RETIREMENT- Severity:
RECOVERABLE- Scope:
DEVICE- Legacy Xid:
156
A TPC was retired and replaced with a TPC from a different GPC.
BANK_REMAPPING_PENDING#
- Code:
0x0016- Category:
RESOURCE_RETIREMENT- Severity:
RECOVERABLE- Scope:
DEVICE- Legacy Xid:
177
A new DRAM bank was marked for remapping. The remap activates on the next GPU reset.
GPU-RC#
CHANNEL_RESET#
- Code:
0x0001- Category:
UNCLASSIFIED_ERROR- Severity:
RECOVERABLE- Scope:
ENGINE- Legacy Xid:
Varies
Robust Channel recovery reset a channel or engine. The legacy Xid context carries the originating robust-channel Xid; the event context carries the root-cause Xid, faulting engine type, channel, TSG, runlist, and instance block.