Analyzing GPU Operational Events#

Warning

GPU Operational Events are a preview feature and remain under active development. Category values, module signatures, module event codes, severities, context payloads, and the set of published events may change before the interface is declared stable. Treat the identifiers in this document as valid for the driver release documented here, and re-validate consumers against each driver release. Legacy Xid output remains the stable contract for production monitoring during the preview period.

Overview#

A GPU Operational Event (GOE) is a structured, machine-readable record that the NVIDIA GPU driver publishes whenever it detects a condition worth reporting to a system operator: an error, a recovery action, a resource retirement, or a notable state change. Each event carries an identity, a severity, an impact scope, and a list of typed context payloads that describe the specific occurrence.

GPU Operational Events give system software a programmatic interface for the same conditions that the driver reports as Xid messages in the kernel log. Each event carries its identity in discrete numeric fields and provides higher attribution resolution than legacy Xid output.

Events that have a legacy Xid equivalent continue to produce their existing NVRM: Xid (PCI:...) kernel log lines unchanged. The structured event and the legacy Xid line describe the same occurrence.

For the actions to take in response to a condition, see the Xid Errors documentation. Each event in Published Events lists its legacy Xid, which selects the corresponding entry in the Xid catalog and its recommended immediate and investigatory actions.

Availability#

GPU Operational Events are available on Linux through NVML, on Turing and later GPUs.

Event Identity#

Event identity is consistent across all GPU Operational Event output formats, including structured records, CPER, and nvidia-smi event-log. Every GPU Operational Event is identified by a three-part tuple. A consumer uses the full tuple to recognize a specific event; individual parts are reused across events.

Field

Type

Description

Module Signature

Character array

NUL-terminated ASCII string naming the source module, for example GPU-GSP. Namespaces the event code.

Category

16-bit integer

Class of condition, from the category registry in Event Categories. Shared across all modules.

Module Event Code

16-bit integer

Specific event within the reporting module. Together with the module signature and category, selects an entry in Published Events.

Each event additionally carries the fields below.

Field

Description

Severity

Impact relative to the event scope. See Event Severities.

Log Level

Urgency of the report, orthogonal to severity. See Event Log Levels.

Scope

What the event affects. See Event Scopes.

Originator

Which driver or firmware component originated the event.

Device UUID

UUID of the GPU that reported the event.

Instance ID

Originator-scoped sequence number.

Timestamp

Microsecond wall clock at event creation.

Trace ID

Group identifier shared by related events that describe one incident.

Group Cursor

64-bit value that identifies the event group. Also maps to the CPER Error Record Header recordId for the retained CPER record.

Group Position

Index of this event within its group.

Attributes

Per-event flags that describe properties of the event.

Group Attributes

Group-level flags that describe properties of the event group.

Attribution

Finer-grained location of the event within the device. GPU Operational Events provide higher attribution resolution than legacy Xid messages.

Event records include grouping fields for correlating related events when the driver reports multiple events for one incident. In this release, published events are delivered individually. The trace ID identifies the incident. See GPU Operational Event Formats for how these fields appear in each published format.

Monitoring GPU Operational Events#

The driver retains recent event groups in an in-memory Operational Event log and publishes them through the interfaces below. All interfaces describe the same events.

The log holds up to 64 event groups and evicts by severity. When the log is full, the driver evicts the oldest group at the lowest resident severity, and does so when the incoming group is at least as severe. When every resident group is more severe than the incoming group, the driver drops the incoming group and counts it as an overflow. Retained history therefore favors the most significant events over the most recent ones, so a long-running consumer should subscribe for live delivery and treat history replay as a catch-up mechanism.

Command Line#

nvidia-smi event-log prints the events that have occurred since driver load in a human-readable form. It reports events for every GPU visible to NVML, skipping any GPU it cannot query. event-log accepts -h / --help. Resolve the numeric values it prints through Published Events. Three properties of the decoder in this release are worth noting when you do:

  • Module event codes currently print as Unknown alongside the numeric value. The numeric value is correct and selects the catalog entry.

  • Contexts other than Legacy Xid print their numeric type and no payload. Legacy Xid contexts print the Xid code.

  • Severity reports the CPER severity of the record rather than the event’s own severity, so a GPU-fatal event appears here as Recoverable. See UEFI CPER for the mapping, and read the severity in a structured record for GPU-relative impact.

nvidia-smi cper prints the same history as Base64-encoded raw CPER records, one record per Base64 block, for forwarding to an external collector. cper accepts -h / --help and --no-wrap. Output wraps at 76 columns by default; --no-wrap emits each CPER record on a single line.

NVML#

NVML exposes GPU Operational Events through two interfaces. See the NVML API Reference Guide for full function prototypes.

Structured event subscription. Register a GPU with nvmlEventSetRegisterGpuOperationalEvents_v1, then retrieve events with nvmlEventSetWait_v3. Records with dataType set to NVML_EVENT_DATA_TYPE_GPU_OPERATIONAL_EVENT carry the structured event. Read attached context payloads with nvmlEventSetGetContextCount_v1, nvmlEventSetGetContextInfo_v1, and nvmlEventSetGetContextData_v1. Decode a legacy Xid context with nvmlEventSetGetGpuOperationalEventContextLegacyXid_v1.

The same event set can subscribe to legacy Xid events and GPU Operational Events. When a GPU Operational Event carries a Legacy Xid context, use that Xid value as a correlation hint to the matching Xid event, kernel log line, and catalog entry.

Retained CPER history. nvmlSystemGetCPER_v1 returns retained CPER records as raw bytes. This is the interface that nvidia-smi cper uses.

Kernel Log#

Events that carry a legacy Xid context emit their traditional kernel log line:

NVRM: Xid (PCI:0000:41:00 GPU-I:03): 63, Row Remapper: New row (0x00000000fbc12000) marked for remapping, reset gpu to activate.

This format is unchanged from previous releases and remains the stable contract for existing log parsers. Events without a legacy Xid equivalent are visible through the structured, CPER, and out-of-band interfaces.

Out-of-Band#

On platforms that support out-of-band reporting, the driver forwards the same rendered CPER record to the BMC, so a management controller observes the event without the host driver stack.

The category registry and the full list of published events for this driver release are in GPU Operational Event Catalog.

Interpreting an Event#

Determine the Reporting Module and Event Code#

Read the module signature and module event code from the structured record, or Source Module and Module Event Code from nvidia-smi event-log. Together with the category, this tuple selects an entry in Published Events, where events are grouped by module signature.

Determine Impact from the Category and Severity#

The category identifies the class of condition and the severity identifies impact relative to the event’s scope. Use scope to establish what the severity applies to: a fatal event at individual GPU engine scope reports that one engine is down, while a fatal event at device scope reports that the GPU is down.

Read the Context Payloads#

Contexts carry occurrence-specific data. In this release, the Legacy Xid context is the one context that NVML and nvidia-smi decode into named fields: the Xid code and its formatted message. A consumer reads the remaining contexts as typed raw bytes and can skip any type it does not recognize using the size field.

Determine the Actions to Take#

For an event that lists a legacy Xid in Published Events, look the Xid up in the Xid Errors catalog, which supplies the immediate and investigatory actions for that condition. Events that report a pending repair also carry the recovery action needed to activate the repair in their context payload.