> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nixl/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nixl/_mcp/server.

# Telemetry Guide

> Understanding NIXL's telemetry system -- event types, metrics, configuration, and monitoring with Prometheus.

## Overview

NIXL's telemetry system collects performance metrics and transfer events for monitoring and debugging. Telemetry is disabled by default and must be enabled via environment variables or agent configuration.

The system supports two consumption patterns:

- **Shared memory cyclic buffer** -- Events are written to a memory-mapped file and read by a separate telemetry reader process. This is the default exporter when `NIXL_TELEMETRY_DIR` is set.
- **Prometheus exporter** -- Events are aggregated and exposed as Prometheus-compatible metrics on an HTTP endpoint.

Only one telemetry exporter plug-in can be loaded per NIXL agent instance.

## Architecture

The telemetry system consists of three layers:

1. **Telemetry Collection** -- Built into the core library, intercepts agent operations (memory registration, transfers, metadata exchange) and generates events with microsecond-precision timestamps.

2. **Event Buffer** -- A cyclic (ring) buffer stores events. Buffer size is configurable via `NIXL_TELEMETRY_BUFFER_SIZE` (default: 4096 events). When the buffer is full, the oldest events are overwritten silently -- the current design allows telemetry loss under high event rates.

3. **Exporter Plugins** -- Flush events from the buffer to an external consumer at configurable intervals. The built-in shared memory buffer exporter uses a memory-mapped file that can be read by external telemetry reader applications. The Prometheus exporter aggregates events into metrics exposed on an HTTP endpoint.

<Note>
The shared memory buffer exporter must be statically linked (built-in module). Other exporters such as the Prometheus exporter are loaded as dynamic plug-ins.
</Note>

### Event Structure

Each telemetry event contains four fields:

| Field | Type | Description |
|-------|------|-------------|
| Timestamp | `uint64` (microseconds) | Microsecond-precision timestamp captured after the operation completes |
| Category | Enum | Event category for filtering and aggregation |
| Event Name | String (max 32 chars) | Descriptive identifier for the specific event |
| Value | `uint64` | Numeric value associated with the event (bytes, count, microseconds) |

<Note>
Timestamps are recorded after the operation completes, not when it starts. Event categories are not a separate on-wire field. To limit exported metrics, use `NIXL_TELEMETRY_ENABLED_METRICS` with a comma-separated fnmatch allowlist of event names.
</Note>

## Event Types

The telemetry system defines eight event categories:

| Category | Description |
|----------|-------------|
| `NIXL_TELEMETRY_MEMORY` | Memory operations -- registration, deregistration, allocation |
| `NIXL_TELEMETRY_TRANSFER` | Data transfer operations -- bytes transmitted/received, request counts |
| `NIXL_TELEMETRY_CONNECTION` | Connection management -- connect and disconnect events |
| `NIXL_TELEMETRY_BACKEND` | Backend-specific operations -- initialization, configuration |
| `NIXL_TELEMETRY_ERROR` | Error events -- error counts by type |
| `NIXL_TELEMETRY_PERFORMANCE` | Performance metrics -- transaction times, latency measurements |
| `NIXL_TELEMETRY_SYSTEM` | System-level events -- process start/stop, resource usage |
| `NIXL_TELEMETRY_CUSTOM` | Custom/user-defined events -- application-specific metrics |

<Note>
Some telemetry categories have no predefined events yet and exist for extensibility. Backend plug-ins can define custom events within any category.
</Note>

## Metrics

The following table lists all built-in telemetry metrics and events generated by NIXL:

| Event Name | Category | Unit | Description |
|------------|----------|------|-------------|
| `agent_memory_registered` | `NIXL_TELEMETRY_MEMORY` | bytes | Registered memory size per registration API call |
| `agent_memory_deregistered` | `NIXL_TELEMETRY_MEMORY` | bytes | Bytes of memory deregistered per API call |
| `agent_tx_bytes` | `NIXL_TELEMETRY_TRANSFER` | bytes | Bytes transmitted by the agent per TX request |
| `agent_rx_bytes` | `NIXL_TELEMETRY_TRANSFER` | bytes | Bytes received by the agent per RX request |
| `agent_tx_requests_num` | `NIXL_TELEMETRY_TRANSFER` | count | Number of transmit requests sent by the agent |
| `agent_rx_requests_num` | `NIXL_TELEMETRY_TRANSFER` | count | Number of receive requests processed by the agent |
| `agent_xfer_time` | `NIXL_TELEMETRY_PERFORMANCE` | microseconds | Transfer time from start to complete (per request) |
| `agent_xfer_post_time` | `NIXL_TELEMETRY_PERFORMANCE` | microseconds | Time from start to posting to backend (per request) |
| Backend-specific events | `NIXL_TELEMETRY_BACKEND` | varies | Dynamic events generated by backend implementations |
| Error status strings | `NIXL_TELEMETRY_ERROR` | count | Error occurrences by status type |
| `agent_telemetry_events_dropped` | `NIXL_TELEMETRY_ERROR` | count | Events dropped at the producer-side staging queue (reported per flush) |

The shared memory buffer exporter stores raw per-event data without aggregation. Each event is recorded individually, preserving the full time-series for offline analysis.

## Enabling Telemetry

Telemetry is controlled by environment variables set before agent initialization:

| Variable | Description | Default |
|----------|-------------|---------|
| `NIXL_TELEMETRY_ENABLE` | Enable telemetry collection | `false` |
| `NIXL_TELEMETRY_ENABLED_METRICS` | Optional comma-separated fnmatch allowlist of event names; unset or empty exports all events | None |
| `NIXL_TELEMETRY_BUFFER_SIZE` | Number of events in the cyclic buffer | `4096` |
| `NIXL_TELEMETRY_RUN_INTERVAL` | Exporter flush interval (milliseconds) | `100` |
| `NIXL_TELEMETRY_EXPORTER` | Name of the exporter plug-in to load | None |
| `NIXL_TELEMETRY_DIR` | Directory for shared memory telemetry files | None |

<Tip>
For the complete list of telemetry-related environment variables including Prometheus settings, see the [Environment Variables](/nixl/resources/environment-variables) page.
</Tip>

`NIXL_TELEMETRY_ENABLE` accepts `y`, `yes`, `on`, `true`, `enable`, or `1` (case-insensitive). Any other value or absence disables telemetry.

**Behavior rules:**

- If telemetry is disabled via the environment but enabled programmatically (`capture_telemetry` in agent config), telemetry is captured internally but not exported. This may incur a performance penalty.
- If telemetry is enabled but no exporter is set and `NIXL_TELEMETRY_DIR` is unset, no telemetry file is generated and `NIXL_TELEMETRY_RUN_INTERVAL` is unused.
- If telemetry is enabled and `NIXL_TELEMETRY_DIR` is set, the shared memory buffer exporter writes events to a file in that directory.

```bash
# Enable telemetry with shared memory buffer exporter
export NIXL_TELEMETRY_ENABLE=true
export NIXL_TELEMETRY_DIR=/tmp/nixl_telemetry
export NIXL_TELEMETRY_BUFFER_SIZE=8192

# Run your NIXL application
./my_nixl_app
```

## Telemetry API

### getXferTelemetry (C++)

Retrieve per-transfer telemetry for a completed transfer request. Requires `captureTelemetry = true` in `nixlAgentConfig`.

| Parameter | Type | Description |
|-----------|------|-------------|
| `req_hndl` | `const nixlXferReqH*` | Transfer request handle |
| `telemetry` | `nixl_xfer_telem_t&` | [out] Populated telemetry data |
| **Returns** | `nixl_status_t` | `NIXL_SUCCESS` or `NIXL_ERR_NO_TELEMETRY` if capture is not enabled |

The returned `nixl_xfer_telem_t` structure contains:
- `startTime` -- Timestamp when the transfer was initiated
- `postDuration` -- Time from start to posting to the backend
- `xferDuration` -- Total transfer time from start to completion
- `totalBytes` -- Total bytes transferred
- `descCount` -- Number of descriptors in the transfer

```cpp
nixl_xfer_telem_t telem;
if (agent.getXferTelemetry(req, telem) == NIXL_SUCCESS) {
    std::cout << "Transfer took " << telem.xferDuration.count() << " us" << std::endl;
    std::cout << "Total bytes: " << telem.totalBytes << std::endl;
}
```

<Tip>
For the complete C++ API reference including transfer and telemetry methods, see the C++ API Reference.
</Tip>

### get_xfer_telemetry (Python)

Python equivalent of the C++ method. Requires `capture_telemetry=True` in `nixl_agent_config`.

```python
telem = agent.get_xfer_telemetry(xfer_handle)
print(f"Transfer duration: {telem.xferDuration} us")
print(f"Total bytes: {telem.totalBytes}")
```

The returned object has the same fields as the C++ version: `startTime`, `postDuration`, `xferDuration`, `totalBytes`, and `descCount`.

<Tip>
For the complete Python API reference including transfer and telemetry methods, see the Python API Reference.
</Tip>

### addTelemetryEvent / getTelemetryEvents (Backend API)

These methods are available on the backend engine base class (`nixlBackendEngine`) and are used by backend plug-in authors to emit custom telemetry events:

- `addTelemetryEvent(event_name, value)` -- Add a custom event to the telemetry buffer
- `getTelemetryEvents()` -- Retrieve recorded events from a backend

These are internal to the backend plug-in interface.

## Telemetry Reader

NIXL provides telemetry reader utilities in both C++ and Python for consuming events from the shared memory cyclic buffer. Below are short snippets showing the core reading loop.

<Tip>
For a complete implementation, see [`examples/python/telemetry_reader.py`](https://github.com/ai-dynamo/nixl/blob/main/examples/python/telemetry_reader.py).
</Tip>

<CodeBlocks>

**`Python`**

```python title="Python"
import time
from nixl_telemetry_reader import SharedRingBuffer

# Open the shared memory telemetry buffer
buffer = SharedRingBuffer("/tmp/nixl_telemetry/agent_name", version=1)
print(f"Buffer capacity: {buffer.get_capacity()} events")

# Read events in a loop
while True:
    event = buffer.pop()
    if event:
        name = event.event_name.decode("utf-8").rstrip("\x00")
        print(f"[{event.timestamp_us}] {name}: {event.value}")
    else:
        time.sleep(0.5)  # No events available, wait briefly
```

**`C++`**

```cpp title="C++"
#include "common/cyclic_buffer.h"
#include "telemetry_event.h"

// Open the shared memory telemetry buffer
sharedRingBuffer<nixlTelemetryEvent> buffer(telemetry_path, false, TELEMETRY_VERSION);
std::cout << "Buffer capacity: " << buffer.capacity() << " events" << std::endl;

// Read events in a loop
nixlTelemetryEvent event;
while (g_running) {
    if (buffer.pop(event)) {
        std::cout << "Event: " << event.eventName_ << std::endl;
        std::cout << "Value: " << event.value_ << std::endl;
    } else {
        std::this_thread::sleep_for(std::chrono::milliseconds(500));
    }
}
```
</CodeBlocks>

**Running the readers:**

```bash
# C++ telemetry reader
./builddir/examples/cpp/telemetry_reader /tmp/nixl_telemetry/agent_name

# Python telemetry reader
python3 examples/python/telemetry_reader.py --telemetry_path /tmp/nixl_telemetry/agent_name
```

**Example output:**

```
=== NIXL Telemetry Event ===
Timestamp: 2025-01-15 14:30:25.123456
Category: TRANSFER
Event: agent_tx_bytes
Value: 1048576
===========================

=== NIXL Telemetry Event ===
Timestamp: 2025-01-15 14:30:25.124567
Category: MEMORY
Event: agent_memory_registered
Value: 4096
===========================
```

## Prometheus Integration

NIXL includes an experimental Prometheus-compatible telemetry exporter that aggregates events into metrics and exposes them on an HTTP endpoint.

### Setup

To enable Prometheus export, set the following environment variables:

```bash
export NIXL_TELEMETRY_ENABLE=true
export NIXL_TELEMETRY_EXPORTER=prometheus
export NIXL_TELEMETRY_PROMETHEUS_PORT=9090     # Default port
export NIXL_TELEMETRY_PROMETHEUS_LOCAL=false    # Bind to all interfaces
```

| Variable | Default | Description |
|----------|---------|-------------|
| `NIXL_TELEMETRY_PROMETHEUS_PORT` | `9090` | HTTP listen port for the metrics endpoint |
| `NIXL_TELEMETRY_PROMETHEUS_LOCAL` | `false` | When `true`, binds to `127.0.0.1` only; when `false`, binds to `0.0.0.0` |

### Scraping Metrics

Once the exporter is running, configure your Prometheus instance to scrape the NIXL metrics endpoint:

```yaml
# prometheus.yml
scrape_configs:
  - job_name: 'nixl'
    static_configs:
      - targets: ['localhost:9090']
    scrape_interval: 5s
```

You can verify the endpoint is serving metrics:

```bash
curl http://localhost:9090/metrics
```

<Warning>
The Prometheus exporter is experimental (beta). It is suitable for development and testing but may change in future releases.
</Warning>

## Custom Telemetry Plug-ins

NIXL supports custom telemetry exporter plug-ins that implement the telemetry export interface. A custom plug-in consumes events from the cyclic buffer and routes them to any monitoring backend -- CSV files, dashboards, or cloud monitoring services.