Initial Incident Report#

When managing GPU system incidents, as in all system incidents, having a well-designed process can help so incidents can be diagnosed more rapidly, and systems can be returned to production faster. At the beginning of any incident, whether detected by the system or reported from a user, try to document the following questions:

  • What was observed about the incident?

  • When was the incident observed?

  • How was it observed?

  • Is this behavior observed on multiple systems or components?

  • If so, how many, and how frequent?

  • Has anything changed with the system, driver, or application behavior recently?

Collecting information regarding these questions to kick-off the debug process is important as it provides the best understanding of the problem and by recording this information can be correlated to other events to gain a better understanding of overall system behavior and health.