Fleet Intelligence - Release Notes

View as Markdown

Version 1.9

New Features:

  • DCGM Error Severity Override [2882]: Added a configurable policy that overrides the effective severity of specific DCGM incidents, for example, downgrading a DCGM code paired with a specific XID from Critical to Warning, without modifying the underlying agent or DCGM evidence.

Changes:

  • API Field Renames [2993]: Renamed several API fields for clarity and consistency with the UI: geoLocation to location, integrityCheck to verificationCheck, integrityCheckReason to verificationCheckReason, integrityCheckExtraInfo to verificationCheckExtraInfo, lastIntegrityCheckTS to lastVerificationCheckTS, and computeZoneGeoLocation to computeZoneLocation. The legacy field names remain available in /v1 as deprecated aliases, and no /v2 API was introduced.
  • Empty and Null Names Sort Last [2879]: Standardized sorting so that empty or null hostname, BMC hostname, and node name values always sort last across the nodes and inventory report APIs.
  • Alert Details Panel Expands Automatically [2846]: The alert details panel now opens already expanded to full screen, removing the extra click previously required to view alert details.

Fixed Bugs:

  • XID Bursts Page Display and Filter Fixes [3070, 3066, 3067, 3082, 3072, 2873]: Fixed several issues on the XID Bursts page: the Data Source field was not hidden when only one agent type was installed and was missing from the filter menu when both agent types were installed, the search placeholder incorrectly read “Search by hostname” and did not search across all machine identifiers, the Machine ID column was shown by default even with only an in-band agent installed, the column order did not follow the design, and column tooltips did not clearly describe each column.
  • Node Group Inventory Pagination Lost at Page Size 100 [2994]: Fixed pagination controls disappearing in the node group inventory view when the page size was set to 100, which blocked aggregate analysis on large clusters.
  • Alerts Page Node Tooltip Missing Detail [2703]: Fixed the Alerts page node tooltip not showing detailed per-node alert information as specified in the design.
  • Notify Alert Rule Dialog Fixes [2531, 2533, 2532]: Fixed the notify alert rule dialog always showing the first rule’s content when editing any rule, showing only the first selected component instead of all selected components, and displaying an incorrectly quoted component name in the pause and activate confirmation message.
  • K8s Agent Enrollment Failure When machine-id Is Not Mounted [3010]: Fixed Kubernetes agent enrollment failing because /etc/machine-id was not mounted into the enroll init container, causing the required machineId field to be rejected during node upsert.

Version 1.8

New Features:

  • XID Bursts: Cloud provider (NCP) users can view groups of related GPU XID messages in the Fleet Intelligence user interface. The page supports filtering, hostname search, configurable columns, and a detailed panel with XID information and recommended actions for data center administrators and tenants. XID Burst data is also available through the Fleet Intelligence API for both NCP and tenant callers.

Changes:

  • Agent Connectivity chart on the Dashboard [2694]: Added an “Agent Connectivity” chart to the Dashboard showing the percentage of machines online over the last 10 days, with a tooltip listing machine detail on hover.
  • Error report filtering by severity [2683, 2684, 2695]: Error reports can now be filtered by error severity, reducing noise when the distribution of severities in a report is uneven.
  • Refreshed CLI report styling [2779]: The CLI-generated report has been visually redesigned.

Fixed Bugs:

  • Alerts table “Sort by Impacted Component” [2856]: Fixed the Level-2 alert table sorting by internal component name instead of the display name shown to users.
  • Missing critical health state capture for HGX firmware components [2776]: Fixed an issue where the out-of-band collector did not capture StandbyOffline/Critical health states reported for certain HGX GPU firmware components.

Version 1.7

New Features:

  • Redesigned alerts experience with full-screen view and comprehensive filters [2397, 2482, 2605, 2616, 2625, 2627, 2628, 2629, 2630]: The alerts experience has been completely redesigned. The alert panel now opens in a dedicated full-screen view with separate tabs for Active Alerts and Historical Alerts. Alerts can be filtered by severity, status, component type, GPU type, node group, compute zone, and machine hostname. Per-machine alert counts are broken down by severity (critical/warning), and machines can be sorted by most critical alerts to prioritize triage.
  • Automatically remove offline machines [2099, 2480, 2481]: Fleet Intelligence can now automatically remove machines that have been offline for a configurable period. The feature is opt-in (off by default) and is configured under Settings > Data Retention. An audit trail of all automatically removed machines is retained in the same settings section. Removed machines continue to appear in inventory and error reports.
  • NGC personal API key support [2555]: Users can now authenticate to the Fleet Intelligence external API using their NGC personal API key, in addition to organization service account keys.
  • Auth status endpoint [2601]: Added a /v1/auth/status public API endpoint that returns whether a given API key is valid and authenticated.
  • Security vulnerability grouping in filters and notifications [2608, 2610, 2660]: PSIRT CVE alert components are now grouped as a single “Security Vulnerability” entry in filter views and notification rule creation. A notification rule targeting “Security Vulnerability” matches alerts for all PSIRT CVEs, while per-CVE mute rules continue to apply to individual CVEs only.
  • Meaningful empty state for telemetry metrics [2474, 2574]: Telemetry charts on the Debugging page now display a contextual empty state message (for example, “No error detected”) for metrics that only emit data when an issue is detected, such as InfiniBand link events. This replaces the generic “Nothing Found” label for these metric types.
  • Attestation DOE unavailable error type [2593]: Added a new attestation error type for when the Device Owner Enablement (DOE) protocol is unavailable.
  • Debugging page metrics filtered by node architecture [2521]: The component and metrics dropdown on the Debugging page now queries the node-specific metrics API, showing only metrics supported by the selected node’s hardware architecture.

Changes:

  • Nodes API last-seen timestamp [2638]: The lastSeen field in the nodes API now reflects the agent’s most recent heartbeat while online, or the time offline status was first detected.
  • API returns 499 on client-cancelled requests [2635]: The backend now returns a 499 status code when a client cancels a request, rather than a 5xx error, reducing false alert noise and improving SLO accuracy.
  • Alert retrigger removed [2511]: Removed alert retriggering behavior; alerts are no longer re-triggered after being resolved.
  • Event bucket dot alignment [2549]: The events/buckets API now includes a firstEventTime field, enabling the UI or customer to accurately position event indicators on the debugging timeline.

Fixed Bugs:

  • Debugging page “Application Error” with multi-component selection [2739]: Fixed an intermittent “Application Error” on the Debugging page when applying filters with a multi-component or multi-metric selection.
  • Hover tooltips disappear on mouse movement [2730]: Fixed icon and table cell hover tooltips vanishing on any mouse movement. Long tooltip content — such as full GPU UUID lists or Slack channel addresses — can now be scrolled and copied.
  • Metrics unavailable for 3d/7d on newly enrolled nodes [2543]: Fixed an issue where metrics were not available for 3-day and 7-day time ranges on newly enrolled nodes when shorter ranges (1h/3h) had data.
  • Hidden columns allow deleting unidentifiable machines [2538]: Fixed a bug where hiding all columns in the machine inventory table rendered blank rows that could still be selected and deleted, with no way to identify which machines were affected.

Version 1.6

New Features:

  • Improved alert notification rules [2436, 2437, 2379, 2380]: Reworked alert notification rule configuration. Notify rules can now be scoped by component and node group, and multiple components can be selected within a single rule.
  • Version columns on machine inventory [2154, 2337, 2338]: Added kernel, firmware, driver, and agent version columns to the machine view and inventory report page, and surfaced when an agent is out of date or an update is available.
  • Full data access through the API [2466]: Exposed all Fleet Intelligence data, including the /metrics endpoint, through the public API.
  • Hosted API documentation [2391]: Added a hosted API documentation endpoint for the Fleet Intelligence API.
  • API firewall and rate limiting [2479]: Added a web application firewall (WAF) and basic rate limiting to the public API backend endpoints.
  • Agent reports check interval [2335]: The agent now includes its check (metric scrape) interval in the node upsert request’s agent configuration.

Changes:

  • Persistence-mode issues reported as warnings [2451]: accelerator-nvidia-persistence-mode errors are now reported as warnings rather than errors.
  • “Point of contact” label [2145]: Renamed the “people in charge” label to “point of contact”.
  • Agent enrollment UUID via dmidecode [2475]: Agent enrollment now retrieves the system (dmidecode) UUID using a dedicated Go package.

Fixed Bugs:

  • Agent version shown incorrectly after upgrade [2315]: Fixed the agent version displaying incorrectly on the dashboard after an agent upgrade (for example, from 1.2.1 to 1.3.0).
  • GPU status metric “Nothing Found” regression [2517]: Fixed a regression where the GPU status metric showed “Nothing Found”.
  • Notification preferences API returned 404 [2519]: Fixed the customer API PUT /v1/notification_preferences endpoint returning a 404.
  • Inconsistent mtbiEnabled across customer APIs [2537]: Fixed mtbiEnabled being inconsistent between the List Customers and Get Customer (by ID) APIs.
  • DCGM query timeouts [2422]: Fixed timeout issues when querying DCGM.

Version 1.5

New Features:

  • Ingestion latency metrics for machine info [2330]: Added backend network/platform latency metrics for the machine info ingestion path to support performance monitoring and troubleshooting of agent-to-backend data flow.
  • OpenAPI 3.1.0 spec [2355]: Backend now consumes the OpenAPI 3.1.0 specification, enabling richer schema features and improved client/SDK generation for the Fleet Intelligence API.
  • Node health history on Debugging page [2367]: The node health history component is now also available in the Events Card on the Debugging page, in addition to the machine side panel.
  • Grouped GPU Clock Event Reasons chart [2272]: The GPU Clock Event Reasons telemetry chart on the Debugging page now groups clock event reasons by GPU. Clicking a GPU label updates the data shown in the chart.
  • Report signing [1089, 2320, 2321]: Added backend and UI support for signed reports. Users can request a digitally signed PDF of an inventory report signed with an NVIDIA signing certificate.

Changes:

  • Sort by GPU Utilization removed from node groups [2512]: The option to sort node groups by GPU Utilization has been removed from the Inventory page.
  • “Last seen” label for agent heartbeat [2148]: The “last updated” timestamp on machine inventory entries has been renamed to “last seen” to better reflect that it represents the agent check-in time rather than a data change timestamp.

Fixed Bugs:

  • GPU Used Memory series rendered in identical color [2233]: Fixed the Telemetry chart’s GPU Used Memory series color consistency.
  • Follow-ups for Next.js/React upgrade [2339]: Fixed two accessibility bugs (button-in-button nesting and missing DialogTitle) and cleaned up React 19 patterns (switched to ref-as-prop and tightened empty list keys).
  • fleetint --version warning output [2043]: Fixed two spurious warnings emitted when running fleetint --version.
  • Suggested action missing from health status tooltip [2288]: Fixed the Suggested Action section.

Version 1.4.1 and 1.4.2

Changes:

  • Agent-managed compute zones and node groups [2280, 2281]: Compute zones and node groups are now owned and managed by the Fleet Intelligence agent rather than created manually via the UI or API. The following operations have been removed: creating and renaming compute zones, creating and updating node groups, and assigning nodes to node groups. These relationships are now established automatically by the agent.

Fixed Bugs:

  • XID burst window showing identical start and end times [2262]: Fixed the XID burst analysis window displaying the same timestamp for both start and end time when multiple XID events occurred within the burst.

Version 1.4

New Features:

  • Node health history [1447, 1863]: A new health history panel in the Machine Details sidebar shows a timeline of node health status changes over time. Operators can expand individual status entries to view component-level health detail and quickly identify nodes with recurring issues.
  • XID burst analysis in alert timeline [1238, 2231]: XID burst analyzer results are now surfaced directly in alert timelines for XID component alerts. Each detected burst shows the burst category (e.g., “Off the Bus”, “DRAM(HBM)”), job-disruption status, burst time window, per-XID breakdown with mnemonic and severity, and recommended actions. Burst analysis events are visible to cloud provider (NCP) users. A per-customer allowlist controls which customers receive alert enrichment.
  • Ampere and Ada Lovelace GPU support [2098]: Fleet Intelligence now supports NVIDIA Ampere and Ada Lovelace data center GPUs. Supported additions include A100 (40GB and 80GB, PCIe and SXM4) and Ada Lovelace (L40, L40S), which were previously rejected as unqualified.
  • Updated telemetry chart labels [1747]: Telemetry chart tooltips now display all relevant label key-value pairs in a Grafana-compatible format. Chart legends support scrolling and allow focusing on individual series.
  • Integrity Check UI updates [2080]: The Integrity Check section of the UI has been updated with a new layout and interaction model.
  • Webhook referenceId uses ncaId [2081]: Webhook alert notifications now use ncaId as the referenceId field for consistent customer identification across integrations.
  • Agent liveness via dedicated metric [1951]: Agent liveness detection now uses the dedicated fleetint_agent_up metric emitted by the agent at each export interval, enabling more reliable liveness signaling independent of node metadata updates. Older agents automatically fall back to the previous detection method.

Fixed Bugs:

  • GPU states not shown for newer agent [2286]: Fixed a regression where the gpu_states chart showed no data in the node detail panel for nodes enrolled with newer agent versions.
  • Agent enrollment with partial NVML GPU visibility [2033]: Fixed a bug where the agent failed to export initial machine info when one GPU on a multi-GPU node was not visible to NVML. The agent now exports data for all visible GPUs rather than aborting the entire export.
  • Power Violation Time Y-axis formatting [2029]: Fixed Y-axis labels in the Power Violation Time telemetry chart displaying fractional minute values (e.g., “8.33 min”) instead of clean integer labels.
  • Inconsistent decimal places in GPU Power telemetry [2027]: Fixed inconsistent decimal display in GPU Power Usage and GPU Power Utilization telemetry panels.
  • GPU Memory Utilization tooltip label [2026]: Fixed the telemetry chart tooltip showing “GPU Memory Copy Utilization” instead of “GPU Memory Utilization” for the GPU Utilization (DCGM) component.
  • XID burst analyzer accuracy [2231]: Fixed several accuracy issues in XID burst analysis: the job_disruption flag now uses authoritative catalog lookup instead of a heuristic; burst duration is now correctly computed; open (in-progress) bursts are no longer written to the timeline prematurely; VBIOS version is now correctly read from the node resource schema.

Version 1.3

New Features:

  • Cross-org alert notification rules [1740, 1744, 1745, 1746, 1792, 1831]: Cloud providers (NCPs) can now create and manage their own alert notification rules scoped to tenant resources. NCP and tenant notify rules are fully isolated — each party sees and manages only their own rules. Mute rules remain visible to both parties as before.
  • New events card on Debugging page [1651, 1856, 1868]: The events card on the Debugging page has been redesigned with a new layout and rendering model. Events and telemetry charts now both render only after submitting the filters form, and a new summarization API returns event counts per time bucket to improve readability when there are many events in the selected period.
  • Enhanced chart legend [1893]: Chart legends now handle a large number of entries with improved scrollability and layout.
  • Incident details in alert timeline [1887, 1946]: Incident details from agent state events are now stored with alerts and surfaced in the alert timeline side panel.
  • Improved telemetry chart label display [1692]: Metrics charts now display all relevant labels in the new label format on hover, with improved handling for multiple labels per series.
  • Separate machine info export endpoint [1739]: Machine inventory information (CPU, GPU, OS, driver versions) is now exported via a dedicated endpoint rather than bundled into OTLP telemetry, following standard OpenTelemetry practices. Liveness heartbeats are also now independent of the telemetry export interval.
  • Remove redundant disk info from node side panel [1915, 1936]: Disk used data has been removed from the machine detail side panel since it is already available in the metrics charts.

Fixed Bugs:

  • Excessive decimal places in telemetry panels [1880, 2030]: Fixed telemetry panel values displaying four decimal places; values now display with two decimal places.
  • Duplicate Running PIDs telemetry for OS component [1882, 2028]: Fixed a bug where “Running PIDs” telemetry was erroneously included as the first entry for any component after the OS component had been selected once in the session.
  • XID analyzer parsing error [2001]: Fixed the XID burst analyzer to use the raw kernel message (extra_info.data.raw_kmsg) instead of the processed health-state message, so XID patterns are correctly matched.
  • GB200 GPU duplicate serial numbers [1879]: Fixed duplicate serial numbers reported for GB200 GPUs in the machine details inventory.
  • Page not refreshing after node deletion [1869]: Fixed a bug where the node list did not auto-refresh after a successful node deletion, leaving the deleted node visible and actionable.
  • hasEnrolledMachines missing from API response [1903]: Fixed the /v1/customers API response to include the hasEnrolledMachines field.
  • XID suggested action text [1560, 1442]: Fixed unmapped suggested action codes (INVESTIGATE_SW/USER, REPORT_ISSUE (IF SEEN >1 PER DAY)) displaying as raw codes instead of user-friendly text. Also fixed a typo: “Invetigatory” → “Investigatory”.

Version 1.2

New Features:

  • Fleet Intelligence API: Enable Fleet Intelligence API for customers.
  • Migrate to V2 API [1791]: Migrate from NGC V1 API to NGC V2 API.
  • Enhance agent log [1778]: Add event_id(UUID) to events info in agent message.
  • Surface DCGM version [1775, 1694, 1695, 1490]: Show the DCGM version on the Machine Details page.
  • Enable “debug” button on backend component alert [1766]: The work is done on the backend, surface the button for Agent Connectivity, Firmware Version, etc.
  • User API access [1642, 1636]: Allow users to access the API.
  • Scopes for API keys [1641]: Allow API keys to be scoped to one of 3 levels for API access.
  • Show available storage [1623, 1469]: The storage graphs show usage. Add available storage line to the graphs.
  • Power consumption as % of possible [1457]: Provide GPU power consumption as a % of possible power draw.
  • Add flag to disable local metrics port [1900]: Add a flag to disable the local metrics port.

Fixed Bugs:

  • Clean up error messages [1801]: Fix a bad error message when a customer ID didn’t exist.
  • Clean up tooltips [1773]: Tooltips display the permissions needed if the user doesn’t have the permissions and the action is disabled appropriately.
  • Fix map pull [1762]: Load the map locally to enhance performance.
  • Double soft machine delete [1753]: Fix a bug where if the machine was soft deleted twice, without page refresh, an error was displayed.
  • Fix graphs with divide by zero [1749]: Fix a bug where if the graph had a divide by zero, the graph would not display. Display a warning message instead.
  • Fix agent export failure [1735]: Fix a bug where the agent export would fail due to a driver hang.
  • Remove duplicate tooltips [1734, 1733, 1732]: Remove duplicate tooltips.
  • Improve Alert Detail [1731]: Add line breaks to Component/Status/Reason.
  • No NVIDIA GPU detected [1724]: When the driver is not installed, the agent should still detect the GPU and report a driver issue.
  • Better verification checking on URL [1717]: Better checking on the URL passed to the —server-url flag.
  • Security enhancements [1715, 1589]: Prevent access to the agent pod from outside the pod. Only for Helm-based install. Set .fleetint files to 0700.
  • Enhance agent docs [1713]: Add —retention and —compact flags to the agent docs.
  • No events in alert timeline [1902]: No events in the alert timeline when an alert is manually muted or unmuted.
  • RTX alert suppression [1876]: Suppress IMEX alerts for RTX cards since they do not apply.
  • Security hardening [1836, 1837, 1838, 1839, 1840, 1841, 1842]: Fix several HTTPS checks and add several security checks.

Version 1.1

New Features:

  • Mute Alert: Added the ability to create and manage rules to mute alerts.

  • Notify Alert: Added the ability to create and manage notify alert rules. Supported channels are:

    • Email
    • Slack (via email)
    • Webhook
  • CVE checking: Added checks for nodes against the NVIDIA CVE database. Added as Metrics for the ability to Mute.

  • New Metrics: Added the following new metrics:

    • dcgm_fi_dev_clocks_event_reasons
    • dcgm_fi_dev_fabric_manager_status
    • dcgm_fi_dev_nvlink_count_symbol_ber_float
    • dcgm_fi_dev_nvlink_count_effective_ber_float
  • GPU Status chart: Added a GPU Status chart to the Dashboard Resource Stats, Utilization section.

  • SXID error suggested actions: Suggestions in events and error reports.

  • Summarize events: Displays a summary of the same events on a node.

  • Update Agent advice: Agent install advice now displays the latest version of the agent.

  • Display larger charts: Added a button to display larger charts in the detail and debugging pages.

  • Search by hostname: Added a search by hostname to the Inventory page.

  • Allow component telemetry: Allow component telemetry to be selected in the debugging pages.

  • Agent Precheck script: Added a precheck script to the agent install.

  • Agent should not enroll without GPUs or with incorrect GPUs: The Agent should not enroll any nodes without GPUs or with incorrect GPUs.

Fixed Bugs:

  • XID display bug: Fixed the accelerator-nvidia-error-sxid has no display name bug.
  • GPU reported up wrongly bug: Fixed GPU State incorrectly remained “up” in the face of XID 94 and 95.
  • Compute Zone View stuck bug: Fixed Compute Zone View stuck in loading state on Inventory page.
  • NVLink BER metrics bug: Fixed NVLink BER metrics showing all zeros in debugging page.
  • Error report dialog hang: Fixed the error report modal dialog hanging after clicking Generate.
  • Metric X-axis truncation bug: Fixed the metric X-axis truncation bug.
  • Machine index: Fixed the GPU index on machine details to match GPU Chart tooltip index.
  • Agent Liveness Check Improvement: Fixed the agent liveness check improvement.
  • Panel Coordination bug: Event/alert from node detail panel should carry to the Debug screen.