Analyzing GPU Operational Events#
Warning
GPU Operational Events are a preview feature and remain under active development. Category values, module signatures, module event codes, severities, context payloads, and the set of published events may change before the interface is declared stable. Treat the identifiers in this document as valid for the driver release documented here, and re-validate consumers against each driver release. Legacy Xid output remains the stable contract for production monitoring during the preview period.
Overview#
A GPU Operational Event (GOE) is a structured, machine-readable record that the NVIDIA GPU driver publishes whenever it detects a condition worth reporting to a system operator: an error, a recovery action, a resource retirement, or a notable state change. Each event carries an identity, a severity, an impact scope, and a list of typed context payloads that describe the specific occurrence.
GPU Operational Events give system software a programmatic interface for the same conditions that the driver reports as Xid messages in the kernel log. Each event carries its identity in discrete numeric fields and provides higher attribution resolution than legacy Xid output.
Events that have a legacy Xid equivalent continue to produce their existing
NVRM: Xid (PCI:...) kernel log lines unchanged. The structured event and the legacy Xid line
describe the same occurrence.
For the actions to take in response to a condition, see the Xid Errors documentation. Each event in Published Events lists its legacy Xid, which selects the corresponding entry in the Xid catalog and its recommended immediate and investigatory actions.
Availability#
GPU Operational Events are available on Linux through NVML, on Turing and later GPUs.
Event Identity#
Event identity is consistent across all GPU Operational Event output formats, including
structured records, CPER, and nvidia-smi event-log. Every GPU
Operational Event is identified by a three-part tuple. A consumer uses the full tuple to
recognize a specific event; individual parts are reused across events.
Field |
Type |
Description |
|---|---|---|
Module Signature |
Character array |
NUL-terminated ASCII string naming the source module, for example |
Category |
16-bit integer |
Class of condition, from the category registry in Event Categories. Shared across all modules. |
Module Event Code |
16-bit integer |
Specific event within the reporting module. Together with the module signature and category, selects an entry in Published Events. |
Each event additionally carries the fields below.
Field |
Description |
|---|---|
Severity |
Impact relative to the event scope. See Event Severities. |
Log Level |
Urgency of the report, orthogonal to severity. See Event Log Levels. |
Scope |
What the event affects. See Event Scopes. |
Originator |
Which driver or firmware component originated the event. |
Device UUID |
UUID of the GPU that reported the event. |
Instance ID |
Originator-scoped sequence number. |
Timestamp |
Microsecond wall clock at event creation. |
Trace ID |
Group identifier shared by related events that describe one incident. |
Group Cursor |
64-bit value that identifies the event group. Also maps to the CPER Error Record
Header |
Group Position |
Index of this event within its group. |
Attributes |
Per-event flags that describe properties of the event. |
Group Attributes |
Group-level flags that describe properties of the event group. |
Attribution |
Finer-grained location of the event within the device. GPU Operational Events provide higher attribution resolution than legacy Xid messages. |
Event records include grouping fields for correlating related events when the driver reports multiple events for one incident. In this release, published events are delivered individually. The trace ID identifies the incident. See GPU Operational Event Formats for how these fields appear in each published format.
Monitoring GPU Operational Events#
The driver retains recent event groups in an in-memory Operational Event log and publishes them through the interfaces below. All interfaces describe the same events.
The log holds up to 64 event groups and evicts by severity. When the log is full, the driver evicts the oldest group at the lowest resident severity, and does so when the incoming group is at least as severe. When every resident group is more severe than the incoming group, the driver drops the incoming group and counts it as an overflow. Retained history therefore favors the most significant events over the most recent ones, so a long-running consumer should subscribe for live delivery and treat history replay as a catch-up mechanism.
Command Line#
nvidia-smi event-log prints the events that have occurred since driver load in a
human-readable form. It reports events for every GPU visible to NVML, skipping any GPU it cannot
query. event-log accepts -h / --help. Resolve the numeric values it prints through
Published Events. Three properties of the decoder in this release are worth noting when you do:
Module event codes currently print as
Unknownalongside the numeric value. The numeric value is correct and selects the catalog entry.Contexts other than Legacy Xid print their numeric type and no payload. Legacy Xid contexts print the Xid code.
Severityreports the CPER severity of the record rather than the event’s own severity, so a GPU-fatal event appears here asRecoverable. See UEFI CPER for the mapping, and read the severity in a structured record for GPU-relative impact.
nvidia-smi cper prints the same history as Base64-encoded raw CPER records, one record per
Base64 block, for forwarding to an external collector. cper accepts -h / --help and
--no-wrap. Output wraps at 76 columns by default; --no-wrap emits each CPER record on a
single line.
NVML#
NVML exposes GPU Operational Events through two interfaces. See the NVML API Reference Guide for full function prototypes.
Structured event subscription. Register a GPU with
nvmlEventSetRegisterGpuOperationalEvents_v1, then retrieve events with
nvmlEventSetWait_v3. Records with dataType set to
NVML_EVENT_DATA_TYPE_GPU_OPERATIONAL_EVENT carry the structured event. Read attached context
payloads with nvmlEventSetGetContextCount_v1, nvmlEventSetGetContextInfo_v1, and
nvmlEventSetGetContextData_v1. Decode a legacy Xid context with
nvmlEventSetGetGpuOperationalEventContextLegacyXid_v1.
The same event set can subscribe to legacy Xid events and GPU Operational Events. When a GPU Operational Event carries a Legacy Xid context, use that Xid value as a correlation hint to the matching Xid event, kernel log line, and catalog entry.
Retained CPER history. nvmlSystemGetCPER_v1 returns retained CPER records as raw bytes.
This is the interface that nvidia-smi cper uses.
Kernel Log#
Events that carry a legacy Xid context emit their traditional kernel log line:
NVRM: Xid (PCI:0000:41:00 GPU-I:03): 63, Row Remapper: New row (0x00000000fbc12000) marked for remapping, reset gpu to activate.
This format is unchanged from previous releases and remains the stable contract for existing log parsers. Events without a legacy Xid equivalent are visible through the structured, CPER, and out-of-band interfaces.
Out-of-Band#
On platforms that support out-of-band reporting, the driver forwards the same rendered CPER record to the BMC, so a management controller observes the event without the host driver stack.
The category registry and the full list of published events for this driver release are in GPU Operational Event Catalog.
Interpreting an Event#
Determine the Reporting Module and Event Code#
Read the module signature and module event code from the structured record, or Source Module and
Module Event Code from nvidia-smi event-log. Together with the category, this tuple selects
an entry in Published Events, where events are grouped by module signature.
Determine Impact from the Category and Severity#
The category identifies the class of condition and the severity identifies impact relative to the event’s scope. Use scope to establish what the severity applies to: a fatal event at individual GPU engine scope reports that one engine is down, while a fatal event at device scope reports that the GPU is down.
Read the Context Payloads#
Contexts carry occurrence-specific data. In this release, the Legacy Xid context is the one
context that NVML and nvidia-smi decode into named fields: the Xid code and its formatted
message. A consumer reads the remaining contexts as typed raw bytes and can skip any type it
does not recognize using the size field.
Determine the Actions to Take#
For an event that lists a legacy Xid in Published Events, look the Xid up in the Xid Errors catalog, which supplies the immediate and investigatory actions for that condition. Events that report a pending repair also carry the recovery action needed to activate the repair in their context payload.