RAS Specifications#

Reliability, availability, and serviceability (RAS) features help detect, report, and diagnose hardware and firmware faults. The following tables describe standard system event log (SEL) errors and original equipment manufacturer (OEM) custom records used for event reporting and service analysis.

Standard SEL Error Types#

SEL Error Type records#

SEL Error Type

Record ID

Record Type

TimeStamp

GeneratorId

EvMRevision

SensorType

SensorNumber

EventDirType

OEM Ev Data 1

OEM Ev Data 2

OEM Ev Data 3

Process Error (AMI BIOS)

XXXXh

02h

XXXXXXXXh

0121h

04h

07h

E8h

6Fh

A5h

xxh

xxh

SpareCore (AMI BIOS)

XXXXh

02h

XXXXXXXXh

0121h

04h

07h

E2h

6Fh

A5h

xxh

xxh

IioInternal (AMI BIOS)

XXXXh

02h

XXXXXXXXh

0121h

04h

07h

E3h

6Fh

A5h

xxh

xxh

BAD E1.S

XXXXh

02h

XXXXXXXXh

0121h

04h

0Dh

EAh

6Fh

A1h

xxh

xxh

UPI (AMI BIOS)

XXXXh

02h

XXXXXXXXh

0121h

04h

13h

E0h

6Fh

A5h

xxh

xxh

UPI Fail Over Error

XXXXh

02h

XXXXXXXXh

0121h

04h

13h

E1h

6Fh

A5h

xxh

xxh

IEH (AMI BIOS)

XXXXh

02h

XXXXXXXXh

0121h

04h

13h

E4h

6Fh

Axh

xxh

xxh

LeakyBucket (AMI BIOS)

XXXXh

02h

XXXXXXXXh

0121h

04h

13h

E5h

6Fh

Axh

xxh

xxh

PCIe (PAS 3.0.1.3)

XXXXh

02h

XXXXXXXXh

0121h

04h

13h

E6h

6Fh

Axh

xxh

xxh

M.2 (PAS 3.0.1.5)

XXXXh

02h

XXXXXXXXh

0121h

04h

C6h

C6h

6Fh

Axh

xxh

xxh

POST (EAS 6.2.2.8 Nvidia Post)

XXXXh

02h

XXXXXXXXh

0121h

04h

C7h

C7h

6Fh

xxh

xxh

xxh

Restore Default

XXXXh

02h

XXXXXXXXh

0121h

04h

C7h

C7h

6Fh

A6h

xxh

xxh

Secure Boot Fail

XXXXh

02h

XXXXXXXXh

0121h

04h

C7h

C7h

6Fh

A7h

Reserved

Reserved

Fail to Find Boot Device

XXXXh

02h

XXXXXXXXh

0121h

04h

C7h

C7h

6Fh

ABh

Reserved

Reserved

OEM Debug Mode

XXXXh

02h

XXXXXXXXh

0121h

04h

C7h

C7h

6Fh

AEh

xxh

xxh

RTC Rest (NV)

XXXXh

02h

XXXXXXXXh

0121h

04h

CDh

CDh

6Fh

Reserved

Reserved

Reserved

MCA (PAS 3.0.1.2)

XXXXh

02h

XXXXXXXXh

0121h

04h

CEh

CEh

6Fh

A1h

xxh

xxh

Standard SEL Error Type Descriptions#

Process Error (AMI BIOS)#

In the firmware-first model, SBIOS enables the processor’s enhanced Machine Check Architecture handling. An uncorrectable or fatal machine-check event, or a correctable machine-check interrupt, generates an SMI. SBIOS collects the MCA-bank information, publishes an extended error log for the operating system, and creates a generic Type 2 processor-error SEL record.

SpareCore (AMI BIOS)#

When the processor RAS handler detects a persistent mid-level cache error in an MLC MCA bank, it requests replacement of the failed core by an on-die spare core at the next reset and records the spare-core event in the SEL.

IioInternal (AMI BIOS)#

This event records errors detected by the CPU Integrated I/O module and its PCIe-related functional units. The reported units include VT-d, the CPU-mesh to-IIO interface, DMA and IIO ring logic, and the inbound and outbound traffic controllers.

BAD E1.S#

During POST, SBIOS initializes E1.S NVMe drives connected through the Broadcom- switch downstream PCIe interfaces. It logs this event when drive initialization times out or fails, identity data cannot be read, active namespace enumeration fails, or submission/completion queue creation fails.

UPI (AMI BIOS)#

The UPI event corresponds to KTI/UPI link error reporting. Link CRC protects against transient data errors, link-level retry retransmits corrupted data, and lane isolation identifies a persistently failing lane. SBIOS logs a detected CRC event as correctable or uncorrectable according to whether retry or reset corrected it, including the CPU, UPI port, severity, and boot state.

UPI Fail Over Error#

KTI/UPI lane failover uses dynamic link-width reduction to recover from a hard failure of one or more data lanes. When possible, the link remains active at a narrower width; SBIOS records the CPU, UPI port, and failed lanes in the SEL.

IEH (AMI BIOS)#

The processor uses a unified, hierarchical Integrated Error Handler with a global IEH in UBOX and satellite handlers in IIO components such as PCIe root ports, traffic-switch and ring logic, VT-d, and data-streaming accelerators. The IEH event reports errors from those IIO-stack units and identifies the local handler by bus, device, and function.

LeakyBucket (AMI BIOS)#

The PCIe root-port leaky-bucket mechanism filters expected correctable link errors, groups bursts into a single event, and lets the counter decay according to the programmed acceptable bit-error rate. If the counter exceeds its threshold, the root port attempts recovery by re-equalizing or retraining the link, potentially at a lower speed or width, and logs an SEL event.

PCIe (PAS 3.0.1.3)#

PCIe device errors are collected by the root-complex event collector and propagated through the satellite and global IEH hierarchy. In firmware-first mode the event triggers an SMI; SBIOS reads the PCIe status registers, sends a base and extended error SEL to the BMC, provides an extended log to the operating system, and then clears the error.

M.2 (PAS 3.0.1.5)#

The M.2 boot drives must be present as NVMe SSDs with a valid GPT EFI partition, a VFAT file system, a readable \\EFI\\BOOT\\BOOTX64.EFI loader, and a working simple-file-system protocol. SBIOS logs this error when any of those requirements is not met.

POST (EAS 6.2.2.8 Nvidia Post)#

SBIOS sends POST-start and POST-end events to the BMC and sends a separate boot-complete event. The SEL distinguishes firmware error, POST start, POST end, and boot completion; for applicable records it also carries the last POST code or the selected boot device and instance.

Restore Default#

SBIOS logs this event when a user loads default values from the setup menu. The event data identifies SBIOS as the settings source and indicates a restore to factory defaults.

Secure Boot Fail#

When a boot device cannot be used because Secure Boot authentication fails, SBIOS logs a dedicated SEL rather than the generic boot-device failure event.

Fail to Find Boot Device#

During the Boot Device Selection phase, SBIOS checks the available boot devices. If none can be used to boot the system, it logs this SEL to identify the condition.

OEM Debug Mode#

The platform exposes a BMC-controlled debug or expert mode. SBIOS reads the mode’s enablement state through IPMI and logs an SEL showing whether debug mode is disabled or enabled.

RTC Rest (NV)#

During boot, S3M detects a bad CMOS battery or loss of power and records that condition in the RTC power-status register. SBIOS logs the RTC-reset event as soon as the IPMI transport is available during the early PEI phase.

MCA (PAS 3.0.1.2)#

For machine-check and corrected-machine-check events, SBIOS collects the MCA bank state in firmware-first mode. In addition to the generic processor SEL, it creates an extended MCA SEL containing the CPU and bank identity, error or failed-device information, severity, and whether the event belongs to the current or previous boot.

OEM SEL Records (Types C0h–DFh)#

The six OEM-defined columns correspond to bytes 11-16 of the record as shown in the table.

OEM SEL Record - Type C0h-DFh#

SEL Error Type

Record ID

Record Type

TimeStamp

Manufacturer ID

OEM Defined (Byte 11)

OEM Defined (Byte 12)

OEM Defined (Byte 13)

OEM Defined (Byte 14)

OEM Defined (Byte 15)

OEM Defined (Byte 16)

PCIe Ext-Error (PAS 3.0.1.3)

XXXXh

C0h

XXXXXXXXh

47h/16h/00h

Vendor ID / xxxxh

Device ID / xxxxh

Reserved

PCIEerrorID / xxh

Memory (PAS 3.0.1.1)

XXXXh

C1h

XXXXXXXXh

47h/16h/00h

McBank number

Type17 Handle / xxxxh

Axh

xxh

xxh

Device Boot Fail

XXXXh

C5h

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Reserved

Reserved

Reserved

OEMEvData1

OEMvData2

Memory Persistent Error (0xCA)

XXXXh

CAh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Type17 Handle / xxxxh

OEMEvData1

OEMEvData2

OEMvData3

Memory Transient Error (0xCB)

XXXXh

CBh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Type17 Handle / xxxxh

OEMEvData1

OEMEvData2

OEMvData3

ECS

XXXXh

CDh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Reserved

Reserved

OEMEvData1

OEMEvData2

OEMvData3

DIMM Mapout

XXXXh

CEh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

OEMEvData1

OEMEvData2

OEMvData3

Reserved

Reserved

Reserved

Firmware Update Status

XXXXh

D3h

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Reserved

Reserved

OEMEvData1

OEMEvData2

OEMvData3

Device Missing

XXXXh

DAh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Reserved

Reserved

OEMEvData1

OEMEvData2

OEMvData3

PCIe Link Spd Width dn Train

XXXXh

D9h

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Reserved

OEMEvData1 / link not expected / 00h: link speed / 01h: link width

Expected link speed/width

OEMEvData2 / Actual link speed/link width

OEMvData3 / Bus number

ADDDC Event

XXXXh

DCh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

Reserved

Destination

SourceRank

OEMEvData1

OEMEvData2

OEMvData3

Secure Boot Failure - Device

XXXXh

D8h

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

OEMEvData1

OEMEvData2

OEMvData3

OEMEvData4

Reserved

Reserved

Shutdown Reason (Microsoft)

XXXXh

DDh

XXXXXXXXh

Manufacturer ID / 37h/01h/00h

Sequence Number

Shutdown Reason

Reserved

Bugcheck Parameter (Microsoft)

XXXXh

DEh

XXXXXXXXh

Manufacturer ID / 37h/01h/00h

Sequence Number

Bugcheck Parameter

System Architecture

EWL Fatal Code

XXXXh

DFh

XXXXXXXXh

Manufacturer ID / 47h/16h/00h

OEMEvData1

OEMEvData2

OEMvData3

Reserved

Reserved

Reserved

OEM Error Type Descriptions (Types C0h–DFh)#

PCIe Ext-Error (PAS 3.0.1.3)#

PCIe errors are collected by the root-complex event collector and propagated through the satellite and global IEH hierarchy. In firmware-first mode, SBIOS reads the PCIe status registers, sends SEL records to the BMC, provides an extended error log to the operating system, and then clears the errors. The extended C0h record identifies the device by vendor and device ID and provides a PCIe error ID.

Memory (PAS 3.0.1.1)#

This record covers memory errors that are not reported as ADDDC, PPR, ECS, persistent correctable errors, or transient correctable errors, including uncorrectable memory errors. It identifies the MCA bank and SMBIOS Type 17 handle, along with the error type, memory controller, boot state, socket, channel, and DIMM.

Device Boot Fail#

During the UEFI Boot Device Selection phase, SBIOS loads and dispatches the image selected by the active boot option. If image loading or dispatch from the device fails, SBIOS logs the C5h event with the boot-device type and instance.

Memory Persistent Error (0xCA)#

A persistent memory correctable error cannot be corrected by retry but can be corrected by ECC. The processor tracks it with per-rank counters; when the threshold is reached, SBIOS invokes available advanced RAS mitigation and logs the affected SMBIOS Type 17 handle, rank, controller, DIMM, channel, socket, and boot state.

Memory Transient Error (0xCB)#

Transient memory correctable errors are retry-corrected iMC events recorded in MCA banks. When the configured bank and UBOX thresholds are reached, an SMI notifies the SBIOS RAS handler, which reads the bank status and logs the affected DIMM information.

ECS#

DDR5 Error Check and Scrub reads internal DRAM data and writes corrected data back when an error occurs. After the periodic patrol-scrub-complete SMI, SBIOS reads the DDR5 on-die ECS registers and logs their corrected-error information, including controller, rank, DIMM, channel, socket, and boot state.

DIMM Mapout#

An uncorrectable MCA memory error can otherwise cause repeated system resets during operating-system use. DIMM map-out removes the faulty DIMM after a reset so the system can boot without entering a continuous reset cycle.

Firmware Update Status#

SBIOS supports capsule-based firmware updates for NVMe and TPM devices and logs the update status in a D3h OEM SEL record. The event identifies the target as M.2, U.2, or TPM and carries the associated status data.

Device Missing#

After PCI enumeration, SBIOS verifies the fixed-topology endpoint devices and attempts link recovery when an expected device is absent from PCI configuration space. If the device remains missing after those recovery steps, SBIOS logs its bus, device, and function in the DAh record.

ADDDC Event#

Adaptive Double DRAM Device Correction mitigates persistent memory correctable errors by activating virtual-lockstep regions at bank granularity. When an ADDDC region is activated or changed, SBIOS logs the source and destination rank or bank, event type, channel, DIMM, socket, controller, boot state, and address information.

Secure Boot Failure - Device#

During PCI enumeration, SBIOS loads and dispatches UEFI option ROMs for NIC and HCA functions. If an option ROM cannot be loaded or started because of a downstream-device Secure Boot failure, SBIOS logs a D8h record with the failure type and the device’s segment, bus, device, and function.

Shutdown Reason (Microsoft)#

Microsoft supplies the shutdown reason code read from the Windows reliability registry. The DDh record uses a sequence number to concatenate data from multiple SEL entries and carries the shutdown-reason value.

Bugcheck Parameter (Microsoft)#

When Windows displays a blue screen, Microsoft logs the bugcheck arguments. The DEh record uses a sequence number to combine data from multiple SEL entries and carries the same bugcheck parameter values shown by the BSOD.

EWL Fatal Code#

The Memory Reference Code reports structured warning and error codes during memory initialization. Fatal Enhanced Warning Log events use the DFh OEM record, which stores the EWL major and minor codes and reserves the remaining OEM-defined bytes.

OEM SEL Records (Types E0h–FFh)#

The thirteen OEM-defined columns correspond to bytes 4-16 of the record as shown in the table.

OEM SEL Record - Type E0h-FFh#

SEL Error Type

Record ID

Record Type

OEM Defined

Extension to UMCE Reporting

XXXXh

E0h

Offset

Timestamp

OEMEvData1

MRC EWL

XXXXh

E1h

OEM Defined Data

Bus OOR

XXXXh

E2h

Reserved

Missing Device Power Cycle

XXXXh

E3h

Bus Number

Device Number

Function Number

Count

Reserved

Reserved

Reserved

Reserved

Reserved

Reserved

Reserved

Reserved

Reserved

Global Reset

XXXXh

E4h

OEM Defined Data

retry_rd_err_log

XXXXh

EDh

Reserved

retry_rd_err_log

retry_rd_err_log_set2

retry_rd_err_log_set3

BER Limit Error

XXXXh

FAh

CeStatus

Manufacturer ID 47h/16h/00h

OEMEvData1

OEMEvData2

OEMvData3

Reserved

Reserved

Reserved

eDPC Error

XXXXh

FBh

UceStatus

Error source ID

OEMEvData1

OEMEvData2

OEMvData3

Reserved

Reserved

Reserved

Advisory Non Fatal

XXXXh

FCh

UceStatus

Manufacturer ID 47h/16h/00h

OEMEvData1

OEMEvData2

OEMvData3

Reserved

Reserved

Reserved

Error Counter

XXXXh

FEh

CeStatus

Manufacturer ID 47h/16h/00h

OEMEvData1

OEMEvData2

OEMvData3

Reserved

Reserved

Reserved

OEM Error Type Descriptions (Types E0h–FFh)#

Extension to UMCE Reporting#

The E0h record is the extended Machine Check Event record used by SBIOS to report MCA details that do not fit in the generic event. A sequence of offset values carries a timestamp plus the APIC and bank identifiers, status, address, and miscellaneous MCA register data.

MRC EWL#

The Memory Reference Code detects structured warning and error codes during memory initialization. Its E1h Enhanced Warning Log record stores the major and minor warning codes followed by fields whose meaning depends on the EWL type, such as socket, channel, DIMM, rank, stack, port, or device identifiers.

Bus OOR#

SBIOS assigns fixed root buses during the PEI phase. Exhausting the available bus numbers would reset the system and load the default root-port topology, so SBIOS is designed to remove the out-of-resource condition.

Missing Device Power Cycle#

If an E1.S or HCA endpoint remains missing after PCI bus enumeration, SBIOS uses a power-cycle reset to attempt recovery. SBIOS makes at most three consecutive attempts, logs each attempt with the bus, device, function, and power-cycle count, and then continues POST if the device is still absent.

Global Reset#

Record type E4h is reserved for a global-reset event and defines bytes 4-16 as OEM-defined data.

retry_rd_err_log#

When persistent memory correctable errors may be mitigated by advanced RAS features, SBIOS needs device- and bank-level information from the iMC retry read-error registers. The EDh record stores the three retry_rd_err_log register sets used to provide that information.

BER Limit Error#

For non-CPU-root-port PCIe links, SBIOS measures correctable errors against a bit-error-rate threshold over repeated sampling intervals. When the link exceeds the allowed threshold enough times, SBIOS logs the FAh event with the correctable-error status and device location and can disable further over-limit checking for that device.

eDPC Error#

Enhanced Downstream Port Containment halts traffic below a downstream port after an unmasked uncorrectable error, preventing corrupted traffic from propagating and permitting software-assisted recovery. In firmware-first mode, the trigger signals SBIOS, which identifies the reporting agent and logs its uncorrectable status and device location in the FBh record.

Advisory Non Fatal#

A PCIe Advisory Non-Fatal event lets a detecting agent report an uncorrectable condition as an advisory ERR_COR notification when another component is responsible for handling it. In firmware-first mode SBIOS identifies the reporting device and logs the FCh event; repeated reports are masked after the configured limit.

Error Counter#

CPU IIO root ports count selected link-correctable errors against a programmed threshold. Reaching the threshold generates an internal ERR_COR and an SMI; SBIOS identifies the root port, logs the FEh error-counter record, and disables further correctable-error signaling for the link.