GPU Operational Event Catalog#

Warning

GPU Operational Events are a preview feature and remain under active development. Category values, module signatures, module event codes, severities, context payloads, and the set of published events may change before the interface is declared stable. Treat the identifiers in this document as valid for the driver release documented here, and re-validate consumers against each driver release. Legacy Xid output remains the stable contract for production monitoring during the preview period.

This page lists the shared event category registry and the GPU Operational Events published in this driver release. See Analyzing GPU Operational Events for monitoring interfaces and how to interpret an event, and GPU Operational Event Formats for the formats events are published in.

Event Categories#

The category registry is shared by all reporting modules. A consumer can act on the category alone when it does not recognize a specific module event code.

Category

Value

Indicates

TIMEOUT

0x0001

An operation did not complete within its expected time budget.

MEMORY_INTEGRITY_ERROR

0x0002

A data integrity error in memory or cache, such as ECC, parity, or poison.

INTERCONNECT_ERROR

0x0003

An error on a link between the GPU and another device, such as NVLink, PCIe, or C2C.

CORRUPTION_DETECTED

0x0004

Corrupted driver or device state was detected.

SECURITY_ERROR

0x0005

A security policy violation or an authentication failure.

POWER_ERROR

0x0006

A power delivery or power management fault.

THERMAL_ERROR

0x0007

A thermal limit was reached or a thermal sensor failed.

FIRMWARE_FAULT

0x0008

A fault detected by firmware running on the device.

ENGINE_ERROR

0x0009

An error reported by a GPU engine.

RESOURCE_EXHAUSTED

0x000A

A repair, spare, or tracking resource has been used up.

TRANSLATION_FAULT

0x000B

An address translation or MMU fault.

UNCLASSIFIED_ERROR

0x7FFF

An error that has not yet been assigned a specific category.

INITIALIZATION

0x8000

Component initialization progress or completion.

SHUTDOWN

0x8001

Component shutdown or teardown.

STATE_CHANGED

0x8002

An operational state transition.

CONFIG_CHANGED

0x8003

A configuration change.

RESOURCE_RETIREMENT

0x8004

A resource has been marked for retirement or repair.

RESOURCE_THRESHOLD

0x8005

A resource usage threshold was crossed.

TELEMETRY

0xFFFE

Operational telemetry.

DEBUG

0xFFFF

Engineering and debug data.

This release publishes events in eight of these categories: TIMEOUT, MEMORY_INTEGRITY_ERROR, INTERCONNECT_ERROR, FIRMWARE_FAULT, RESOURCE_EXHAUSTED, UNCLASSIFIED_ERROR, INITIALIZATION, and RESOURCE_RETIREMENT. The remaining values are defined so that consumers can decode events introduced in later releases.

Note

In this release, the nvidia-smi event-log decoder renders category 0x0008 as “Engine Software Error”. The numeric value is correct and corresponds to FIRMWARE_FAULT, so decode categories from the numeric value.

Event Log Levels#

The log level describes the urgency of the report independently from the event severity. Values increase with urgency, so a subscription filter can use “at least this urgent” semantics. A filter value of 0 matches every log level.

Log Level

Value

Indicates

TELEMETRY

10

Operational telemetry.

DIAGNOSTIC

20

Diagnostic information. NVML names this level NVML_GPU_OPERATIONAL_EVENT_LOG_LEVEL_DIAG.

NOTICE

30

A notable condition that does not require immediate action.

WARNING

40

A condition that may require operator attention.

ERROR

50

A condition that requires operator attention.

Event Severities#

The severity identifies the impact of the event relative to its scope. Values increase with severity, so a subscription filter can use “at least this severe” semantics. A filter value of 0 matches every severity.

Severity

Value

Indicates

INFORMATIONAL

10

The event reports state, telemetry, or progress.

CORRECTED

20

The condition was corrected without recovery action by the operator.

RECOVERABLE

30

The condition affected operation, and recovery is possible.

FATAL

40

The condition caused loss of the scoped resource.

When a GPU Operational Event is encoded as CPER, the GOE severity is remapped to a host-scoped CPER severity. See GOE Severity to CPER Severity Mapping.

Event Scopes#

The scope identifies the resource affected by the event.

Scope

Value

Indicates

SYSTEM

0

The whole system.

DEVICE

1

One GPU device.

GPU_INSTANCE

2

One MIG GPU instance.

COMPUTE_INSTANCE

3

One MIG compute instance.

ENGINE

4

One GPU engine.

Published Events#

This release publishes the events listed below. To identify an event you have decoded, find its module signature and module event code in the lookup table, then follow the event name to its entry. Not all Xids emitted by the driver have been fully converted to GPU Operational Events yet; more support will be added in future releases. To determine the actions to take, look up the event’s legacy Xid in the Xid Errors catalog.

A severity listed as FATAL or RECOVERABLE is selected at runtime according to the impact of the reported condition.

Event Lookup#

Module

Code

Event

Legacy Xid

Category

GPU

0x0001

DRIVER_INIT

None

INITIALIZATION

GPU-BUS

0x0002

C2C_CONTAINMENT

179

INTERCONNECT_ERROR

GPU-FSP

0x0001

BOOT_TIMEOUT

143

TIMEOUT

GPU-FSP

0x0002

FUSE_ERROR

143

FIRMWARE_FAULT

GPU-GSP

0x0001

RPC_TIMEOUT

119

TIMEOUT

GPU-GSP

0x0002

LIBOS_HEARTBEAT_TIMEOUT

120

TIMEOUT

GPU-GSP

0x0003

GSP_RM_HEARTBEAT_TIMEOUT

120

TIMEOUT

GPU-GSP

0x0004

FIRMWARE_FAULT

120

FIRMWARE_FAULT

GPU-GSP

0x0005

POISON

140

MEMORY_INTEGRITY_ERROR

GPU-MEMSYS

0x0001

MAINT_OP_ERROR

175

MEMORY_INTEGRITY_ERROR

GPU-NVLINK

0x0001

ALI_TRAINING_FAILURE

136

INTERCONNECT_ERROR

GPU-NVLINK

0x0002

MSE_DEGRADED

150

FIRMWARE_FAULT

GPU-NVLINK

0x0003

MSE_WATCHDOG_TIMEOUT

150

FIRMWARE_FAULT

GPU-NVLINK

0x0004

SW_LINK_DOWN

155

INTERCONNECT_ERROR

GPU-RAS

0x0001

ROW_REMAPPING_PENDING

63

RESOURCE_RETIREMENT

GPU-RAS

0x0002

ROW_REMAPPING_FAILURE_TABLE_FULL

64

RESOURCE_EXHAUSTED

GPU-RAS

0x0003

ROW_REMAPPING_FAILURE_BANK_FULL

64

RESOURCE_EXHAUSTED

GPU-RAS

0x0004

ROW_REMAPPING_FAILURE_INTERNAL_ERROR

64

UNCLASSIFIED_ERROR

GPU-RAS

0x0005

ROW_REMAPPING_FAILURE_RESERVED_ROW

64

RESOURCE_EXHAUSTED

GPU-RAS

0x0006

ROW_REMAPPING_FAILURE_PAGE_OFFLINE_FAILURE

64

UNCLASSIFIED_ERROR

GPU-RAS

0x0007

DRAM_RETIREMENT_FAILURE_INTERNAL_ERROR

64

UNCLASSIFIED_ERROR

GPU-RAS

0x0008

DRAM_RETIREMENT_FAILURE_INFOROM_FULL

64

RESOURCE_EXHAUSTED

GPU-RAS

0x0009

DRAM_RETIREMENT_FAILURE_HW_LIMIT

64

RESOURCE_EXHAUSTED

GPU-RAS

0x000A

DRAM_RETIREMENT_FAILURE_NO_SPARE

64

RESOURCE_EXHAUSTED

GPU-RAS

0x000B

LTS_REPAIR_PENDING

160

RESOURCE_RETIREMENT

GPU-RAS

0x000C

LTS_REPAIR_FAILURE

161

RESOURCE_EXHAUSTED

GPU-RAS

0x000D

MEMORY_CHANNEL_REPAIR_PENDING

160

RESOURCE_RETIREMENT

GPU-RAS

0x000E

MEMORY_CHANNEL_REPAIR_FAILURE

161

RESOURCE_EXHAUSTED

GPU-RAS

0x000F

TPC_REPAIR_PENDING_SAME_GPC

156

RESOURCE_RETIREMENT

GPU-RAS

0x0010

TPC_REPAIR_PENDING_DIFFERENT_GPC

156

RESOURCE_RETIREMENT

GPU-RAS

0x0011

TPC_REPAIR_FAILURE_NO_SPARE

157

RESOURCE_EXHAUSTED

GPU-RAS

0x0012

TPC_REPAIR_FAILURE_NO_SPARE_MIG

157

RESOURCE_EXHAUSTED

GPU-RAS

0x0013

ECC_RESIDUAL_UNCORRECTABLE_ERROR

140

MEMORY_INTEGRITY_ERROR

GPU-RAS

0x0014

DRAM_ECC_INTR_STORM

92

MEMORY_INTEGRITY_ERROR

GPU-RAS

0x0015

SM_ECC_INTR_STORM

92

MEMORY_INTEGRITY_ERROR

GPU-RAS

0x0016

BANK_REMAPPING_PENDING

177

RESOURCE_RETIREMENT

GPU-RC

0x0001

CHANNEL_RESET

Varies

UNCLASSIFIED_ERROR

Event Details#

Entries are grouped by source module signature, then ordered by category value and module event code.

GPU#

DRIVER_INIT#
Code:

0x0001

Category:

INITIALIZATION

Severity:

INFORMATIONAL

Scope:

DEVICE

Legacy Xid:

None

GPU driver initialization completed for the device. Carries driver and device metadata.

GPU-BUS#

C2C_CONTAINMENT#
Code:

0x0002

Category:

INTERCONNECT_ERROR

Severity:

FATAL or RECOVERABLE

Scope:

DEVICE

Legacy Xid:

179

C2C containment was triggered on the chip-to-chip link to the host CPU. Carries the containment code.

GPU-FSP#

BOOT_TIMEOUT#
Code:

0x0001

Category:

TIMEOUT

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

143

The driver timed out polling for FSP boot completion during GPU initialization.

FUSE_ERROR#
Code:

0x0002

Category:

FIRMWARE_FAULT

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

143

The FSP fuse error check failed. Carries the fuse status.

GPU-GSP#

RPC_TIMEOUT#
Code:

0x0001

Category:

TIMEOUT

Severity:

FATAL or RECOVERABLE

Scope:

DEVICE

Legacy Xid:

119

The driver timed out waiting for an RPC response from GSP. Carries the RPC function, name, sequence number, and wait duration.

LIBOS_HEARTBEAT_TIMEOUT#
Code:

0x0002

Category:

TIMEOUT

Severity:

FATAL or RECOVERABLE

Scope:

DEVICE

Legacy Xid:

120

The LibOS heartbeat from the GSP microcontroller stopped.

GSP_RM_HEARTBEAT_TIMEOUT#
Code:

0x0003

Category:

TIMEOUT

Severity:

FATAL or RECOVERABLE

Scope:

DEVICE

Legacy Xid:

120

The GSP-RM heartbeat stopped.

POISON#
Code:

0x0005

Category:

MEMORY_INTEGRITY_ERROR

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

140

GSP consumed poisoned data.

FIRMWARE_FAULT#
Code:

0x0004

Category:

FIRMWARE_FAULT

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

120

GSP firmware reported a fault.

GPU-MEMSYS#

MAINT_OP_ERROR#
Code:

0x0001

Category:

MEMORY_INTEGRITY_ERROR

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

175

A GPU memory subsystem maintenance operation timed out. Carries the operation type.

GPU-RAS#

ECC_RESIDUAL_UNCORRECTABLE_ERROR#
Code:

0x0013

Category:

MEMORY_INTEGRITY_ERROR

Severity:

FATAL or RECOVERABLE

Scope:

DEVICE

Legacy Xid:

140

An uncorrectable ECC error was detected outside the normal handling path. Carries DRAM, LTC, MMU, and PCIe error counts.

DRAM_ECC_INTR_STORM#
Code:

0x0014

Category:

MEMORY_INTEGRITY_ERROR

Severity:

CORRECTED

Scope:

DEVICE

Legacy Xid:

92

Single-bit ECC error interrupts were disabled on a framebuffer partition because of a high error rate. Carries the partition index.

SM_ECC_INTR_STORM#
Code:

0x0015

Category:

MEMORY_INTEGRITY_ERROR

Severity:

CORRECTED

Scope:

DEVICE

Legacy Xid:

92

An SM single-bit ECC interrupt storm was detected.

ROW_REMAPPING_FAILURE_TABLE_FULL#
Code:

0x0002

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

Row remapping failed because the row remapping table is full.

ROW_REMAPPING_FAILURE_BANK_FULL#
Code:

0x0003

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

Row remapping failed because all reserved rows for the bank are already remapped.

ROW_REMAPPING_FAILURE_RESERVED_ROW#
Code:

0x0005

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

Row remapping failed because the target is a reserved row.

DRAM_RETIREMENT_FAILURE_INFOROM_FULL#
Code:

0x0008

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

DRAM retirement failed because the InfoROM object is full.

DRAM_RETIREMENT_FAILURE_HW_LIMIT#
Code:

0x0009

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

DRAM retirement failed because a hardware limit was reached.

DRAM_RETIREMENT_FAILURE_NO_SPARE#
Code:

0x000A

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

DRAM retirement failed because no spare remains for retirement.

LTS_REPAIR_FAILURE#
Code:

0x000C

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

161

L2 slice repair failed because no spare L2 slices remain.

MEMORY_CHANNEL_REPAIR_FAILURE#
Code:

0x000E

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

161

Memory channel repair failed because no spare channels remain.

TPC_REPAIR_FAILURE_NO_SPARE#
Code:

0x0011

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

157

TPC retirement failed because no spare TPCs are available.

TPC_REPAIR_FAILURE_NO_SPARE_MIG#
Code:

0x0012

Category:

RESOURCE_EXHAUSTED

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

157

TPC retirement failed in MIG mode because no spare TPCs remain in the same GPC.

ROW_REMAPPING_FAILURE_INTERNAL_ERROR#
Code:

0x0004

Category:

UNCLASSIFIED_ERROR

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

Row remapping failed due to an internal driver error. Carries an error string.

ROW_REMAPPING_FAILURE_PAGE_OFFLINE_FAILURE#
Code:

0x0006

Category:

UNCLASSIFIED_ERROR

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

The row is already pending remapping. The remap occurs on the next GPU reset.

DRAM_RETIREMENT_FAILURE_INTERNAL_ERROR#
Code:

0x0007

Category:

UNCLASSIFIED_ERROR

Severity:

FATAL

Scope:

DEVICE

Legacy Xid:

64

DRAM retirement failed due to an internal driver error. The context identifies the specific reason.

ROW_REMAPPING_PENDING#
Code:

0x0001

Category:

RESOURCE_RETIREMENT

Severity:

RECOVERABLE

Scope:

DEVICE

Legacy Xid:

63

A new DRAM row was marked for remapping. The remap activates on the next GPU reset.

LTS_REPAIR_PENDING#
Code:

0x000B

Category:

RESOURCE_RETIREMENT

Severity:

RECOVERABLE

Scope:

DEVICE

Legacy Xid:

160

An L2 slice and its pair were marked for repair. Carries the location and the recovery action needed to activate the repair.

MEMORY_CHANNEL_REPAIR_PENDING#
Code:

0x000D

Category:

RESOURCE_RETIREMENT

Severity:

RECOVERABLE

Scope:

DEVICE

Legacy Xid:

160

A memory channel and its pair were marked for repair. Carries the location and the recovery action needed to activate the repair.

TPC_REPAIR_PENDING_SAME_GPC#
Code:

0x000F

Category:

RESOURCE_RETIREMENT

Severity:

RECOVERABLE

Scope:

DEVICE

Legacy Xid:

156

A TPC was retired and replaced with a spare from the same GPC.

TPC_REPAIR_PENDING_DIFFERENT_GPC#
Code:

0x0010

Category:

RESOURCE_RETIREMENT

Severity:

RECOVERABLE

Scope:

DEVICE

Legacy Xid:

156

A TPC was retired and replaced with a TPC from a different GPC.

BANK_REMAPPING_PENDING#
Code:

0x0016

Category:

RESOURCE_RETIREMENT

Severity:

RECOVERABLE

Scope:

DEVICE

Legacy Xid:

177

A new DRAM bank was marked for remapping. The remap activates on the next GPU reset.

GPU-RC#

CHANNEL_RESET#
Code:

0x0001

Category:

UNCLASSIFIED_ERROR

Severity:

RECOVERABLE

Scope:

ENGINE

Legacy Xid:

Varies

Robust Channel recovery reset a channel or engine. The legacy Xid context carries the originating robust-channel Xid; the event context carries the root-cause Xid, faulting engine type, channel, TSG, runlist, and instance block.