Resource Metrics
Resource Metrics
The built-in resource_metrics plugin collects central processing unit (CPU),
memory, process, disk, network, and graphics processing unit (GPU) data. It can
collect on a timer and emit metric marks. A mark is an event that subscribers
can record. Your program can also call collect() for a fresh snapshot. A
snapshot is a set of readings taken at one time.
The plugin is configured through the standard plugin host. Use the same
resource_metrics component in plugins.toml or in a plugin configuration
object. For setup and activation details, refer to
Configure and Initialize Plugins.
Enable the Component
This example enables every category and starts polling every 5 seconds:
Replace / with an absolute path on your system, or use [] when you do not
need filesystem capacity data. The example uses the default values for most
fields; showing them makes each option clear.
The component and polling have separate switches:
components[].enabledactivates the plugin.components.config.polling.enabledstarts or stops periodic collection.components.config.polling.interval_millissets the delay between polls.
Polling is disabled by default. Its default interval is 5,000 milliseconds.
The interval must be greater than zero even when polling is disabled. Each
category is enabled by default. A disabled category is null in the snapshot.
For a ready collection target, the first poll starts when the component
activates. Later polls wait for the configured interval.
Each poll reads current values and running totals. Some values, such as total system memory, can stay the same across many polls. Relay still includes them so each snapshot stands on its own. Relay also checks for devices and processes on each poll. Global scope can query many processes, so collection can take longer. Increase the interval or turn off process input/output (I/O) and GPU process metrics if you do not need them.
Relay keeps the Linux process clock rate and its NVIDIA Management Library (NVML) connection between polls. If NVML is unavailable, Relay tries to find it again after one minute. Relay checks capacity, limits, and device lists on each poll so it can detect changes.
On macOS, GPU sampling runs the system ioreg utility to read driver
diagnostics. Disable the GPU category or its process group if you do not need
those readings.
You can activate the component with polling disabled and call the on-demand collection API. On-demand collection does not require polling to be enabled.
Configuration Fields
All fields use snake_case in plugin configuration. Missing fields use the
defaults below. Unknown fields and invalid values are rejected during
validation.
Measurement Scope
measurement_scope defaults to runtime_default. A scope is the process or
system resources included in a reading. This setting chooses that scope.
The default depends on where the plugin runs. The CLI uses global to measure
system resources visible to Relay. Embedded Rust, Python, and Node.js APIs use
process_tree, rooted at the current process ID.
Set global, application_process, or process_tree to choose a scope. When
you choose process_tree, the CLI measures its launched agent and children.
Embedded APIs measure the tree rooted at their current process.
Polling
These fields control when the plugin collects a snapshot:
Categories
These fields control which measurements each snapshot includes:
Every path in filesystem_paths must be absolute and non-empty. GPU selectors
and network interface names must not be blank. A device index is matched as a
decimal string, such as "0"; not every backend provides an index. Device
identifiers come from the GPU backend and can differ between systems.
Linux AMD and Intel Direct Rendering Manager (DRM) devices do not report an index through this collector. Select them by identifier. macOS registry identifiers can change after a restart. Read a current snapshot before you set a fixed selector.
Turning off disk.enabled disables both disk subgroups. Turning off
process_io keeps filesystem capacity data active. Turning off either GPU
measurement group does not turn off the other group. Network readings are
system-wide in every collection mode. They do not measure an agent process.
Snapshot Shape
Each snapshot has one UTC timestamp, an operating_system, a
measurement_scope, and optional process_sampling metadata. Rust and Python
use snake_case field names. Node.js uses camelCase property names. The
values and meanings are the same.
The snapshot always has cpu, memory, process, disk, gpu, and
network category fields. A disabled category is null. An enabled category
is an object with a stable set of fields. A value that could not be collected
is null, never zero. Each available measurement has only a value and a
unit. This example has CPU enabled and every other category disabled:
Every enabled category contains all of its fields. timestamp is the time
collection finished. operating_system is linux, macos, or windows.
measurement_scope gives the scope that Relay collected: global,
application_process, or process_tree.
Global scope covers resources the operating system exposes to Relay. A
process-tree snapshot includes its root and active children. The scope selects
which process readings to include. System memory, filesystem capacity,
device-wide GPU readings, and network readings describe their wider resources
in every mode. network.measurement_scope is always global.
Some values are running totals. CPU time and disk I/O fields add the lifetime counters of processes visible at collection time. These sums can fall when a process exits. Relay can miss a process that starts and exits between polls.
If a process query shows that a child exited after Relay listed the process tree, Relay skips that child and reports metrics for the remaining live tree. The exited child does not make the process totals unavailable. Relay also excludes that child from the process IDs used for environment and GPU readings. This behavior is the same for polling and on-demand collection in every language.
A process total is still unavailable if Relay cannot read a remaining process,
access is denied, a process ID has been reused, or any sampled process lacks
that counter. If the root exits or its identity changes, Relay returns an
unavailable snapshot. In global scope,
each total adds only the processes that supply that counter. A global total is
null if no process supplies the counter or the sum exceeds its numeric range.
Use process_sampling to check coverage before treating a global total as complete:
visible_processescounts processes listed when collection began, including children that exit before sampling. In process-tree mode, it can therefore be larger thansampled_processeseven when the remaining tree is fully sampled.sampled_processescounts processes whose base query succeeded.field_sampled_processescounts readable values for each process-summed field. Keys use canonical paths, such ascpu.total_time, in every language. A zero count means no process supplied a usable value. Disk throughput keys count processes with a usable change over the sampling interval; they are zero on the first reading. A missing key means Relay did not attempt a process sum or disk-rate calculation for that field. Compare a field’s count withvisible_processesto check its coverage of the listed processes.
The process_sampling metadata is null when Relay did not sample process
counters or could not list processes. Polling marks include three coverage
gauges, each measured in processes:
nemo.relay.resource.process_sampling.visible_countnemo.relay.resource.process_sampling.sampled_countnemo.relay.resource.process_sampling.field_sampled_count, with the canonical field path in thenemo_relay.resource.fieldattribute
If Relay cannot list processes, independent readings such as system memory, filesystem capacity, and device-wide GPU data can still be available.
Arrays have separate empty and unavailable meanings:
limit_eventsis an empty array when no supported limit events are present.disk.filesystemsis an empty array when no paths are configured.network.interfacesis empty when no selected interface is visible.- A GPU subgroup is
nullwhen disabled or collection cannot start. It can be an empty array when collection finds no supported or matching records. - A configured filesystem path stays in
filesystemswhen its capacity query fails; that entry’s capacity measurements arenull.
CPU Category
CPU time uses milliseconds by default. For a process tree or global scope, the time fields add lifetime values from processes visible during collection. These sums can fall when a process exits.
For application-process and process-tree modes, CPU rate uses the change in each process’s CPU time between samples. A newly seen process contributes its current lifetime CPU time. Relay can miss work from processes that exit before the next sample. In global mode, CPU rate comes from system-wide CPU use reported by the operating system. Relay measures the time between samples with a clock that only moves forward, so changes to the system clock do not affect the calculated rate.
The CPU category contains these fields:
consumption_rate is not a percentage. With the default logical_processors
unit, a value of 1.0 means the selected scope used about one logical
processor over the sample interval. One logical processor is one CPU hardware
thread working at full capacity. The same use is 1,000 millicores.
Pressure stall time measures how long work waited for CPU or memory resources. Pressure, throttling, limits, and event counts need operating-system support, so these values can be unavailable.
Very short polling intervals can also leave the global CPU rate null until
the operating system provides another usable reading.
Memory Category
Memory values, including GPU memory, use KiB by default (1 KiB is 1,024 bytes). Collection converts byte and page values to KiB and drops any remainder before applying the selected output unit.
The memory category contains these fields:
The meaning of resident, private, footprint, and environment-accounted memory
depends on the operating system. Refer to the
compatibility tables.
The three system_ fields describe system memory visible to Relay in every
mode. They do not describe the selected process or tree. In global mode,
per-process memory fields are null.
On macOS, system_available follows Apple’s available non-compressed memory
definition. It includes active memory, so it is broader than the vm_stat free
plus inactive pool. system_used and system_available can overlap; their sum
does not necessarily equal system_total.
Process Category
Counts add across the measured process scope unless the description says otherwise. Global counts cover processes visible to the operating system. A process-tree count includes its root and active descendants.
The process category contains these fields:
open_file_descriptor_count is not a count of every Windows kernel handle.
Windows reports those separately as windows_handle_count.
Each limit_events entry has a resource (cpu, memory, or processes),
an event name, and an optional count measurement in events. Current
collectors report CPU throttled, memory high, maximum, or
out_of_memory, and process maximum events where supported.
Disk Category
Disk transfer and filesystem capacity use bytes by default. Process I/O
totals add lifetime counters and can fall when a process exits. Throughput
uses the change between two samples and defaults to bytes per second. A newly
seen process contributes its current lifetime counter to that interval. Relay
can miss transfers from a process that exits before the next sample. Disk rates
in global mode skip processes whose current or previous I/O counter is
unavailable. Check process_sampling.field_sampled_processes for
disk.read_throughput and disk.write_throughput to see how many processes
supplied usable interval values. A rate is unavailable if no process supplies
a usable change, a readable counter decreases, or the sum overflows. Rates for
an application process or process tree require every sampled process to supply
the required current and previous counters.
The operating systems count I/O differently. Linux byte fields count storage traffic, while its operation fields count read and write system calls. macOS reports disk-transfer bytes but not operation counts. Windows reports transfers and operations handled by its I/O manager. Those values can include activity beyond physical disk traffic. Compare samples from the same operating system. The shared unit does not mean the counters measure the same activity.
The disk category contains these fields:
Each filesystems entry contains the configured path and three optional
measurements:
total_capacityis the filesystem’s total capacity.available_capacityis the space available to the calling user.free_capacityis total free space, which can include space reserved for an administrator.
For global scope, process I/O covers all processes visible to the operating
system. For process_tree scope, it covers that tree. Set
filesystem_paths = [] when you only want process I/O. Path capacity is
measured for the filesystem containing each path, so two paths on the same
filesystem can produce repeated capacity values. Throughput is null on the
first sample, after a counter reset, or when Relay cannot read every needed
process counter. Polling and on-demand collection keep separate earlier samples.
Network Category
Network data always describes selected system interfaces. It does not measure
the selected process or its children. The network.measurement_scope field is
global even when the snapshot’s measurement_scope is process_tree.
The category includes system, an aggregate of selected interfaces, and
interfaces, one record per selected interface. Each record has the following
measurements:
The system record has the same measurements. Each interface record also has
its name. Set network.interfaces to names from a current snapshot, or
leave it empty to include all visible interfaces.
An unknown name produces no interface record. If no interfaces match, the
system measurements are null. The first sample has no throughput reading.
A counter reset makes that interval’s throughput null. If a new interface
appears, the system throughput is null until that interface has two readings.
System values add selected interface counters. Traffic can cross both a virtual interface and a physical interface, so the aggregate can count it twice. Interface data totals and process disk totals come from different sources and should not be compared as if they measured the same traffic. On Windows, the network backend omits some disconnected and software interfaces. The aggregate covers the interfaces listed in the snapshot.
GPU Category
The gpu object has two optional groups: device_metrics and
process_metrics. Each record can contain:
Device records describe the whole GPU. Process records describe GPU use by a
process in the measured scope. A record can have one available measurement
and one null measurement if the driver cannot provide both. An empty
devices list selects all supported devices that Relay detects. A selector
can match a reported identifier or an index written as a string.
GPU compute use measures recent activity. NVML reports process activity since
the previous query in the same polling or on-demand series. Linux DRM uses
changes in GPU cycle counters or engine busy time. Apple uses changes in GPU
time reported by its AGX driver. The first sample has no earlier reading, so
process compute use is null. Process memory can still be available. Changing
the collection target clears the earlier reading. GPU process records also
depend on driver access and can be incomplete in global scope.
Configure Units
Every available measurement contains a value and a unit. Relay collects in
consistent base units, then applies your chosen output units. Missing unit
settings keep the defaults. Unit names use snake_case in every API.
The typed APIs group units by what they measure. Each measurement field uses
its matching unit type, so a duration cannot use a capacity or bandwidth unit.
Capacity and transferred data have separate types even though they allow the
same byte units. Rust uses enums. Python and Node.js use string enums. JSON
measurements still contain only a value and a unit, such as
{ "value": 1, "unit": "seconds" }. An unavailable measurement is null.
The following table maps the unit types to their measurements:
This example sets different units for individual measurements:
The following table lists the allowed units for each measurement family:
Each decimal storage unit is 1,000 times larger than the previous unit. Each binary unit is 1,024 times larger than the previous unit. A byte has 8 bits. One logical processor equals 1,000 millicores. A value of 100% equals a fraction of 1.
A unit choice for one field applies to every record of that kind, such as every filesystem or GPU device. System and interface network readings have separate choices. Counts, including packets, errors, processes, and disk operations, keep their named units.
Exact whole-number conversions stay integers. Conversions that produce a
fraction use a decimal number. Rust, Python, and Node.js expose the same
value and unit; Node.js uses bigint for integers above its safe integer
range. If the operating system rounded a reading, changing its unit cannot
restore the lost detail. Measurements that Relay cannot collect or convert
remain null.
Metric marks use a decimal number for a unit that can produce fractions, even
when one reading happens to be whole. This keeps the mark’s number type the
same across polls. Snapshots keep exact integers when the conversion allows
them. The metric event format cannot carry an integer above
9,223,372,036,854,775,807. A snapshot can still contain that value, but
its mark omits the measurement.
If you change a unit after an OpenTelemetry metrics exporter has received that metric, restart the exporter before collecting again. The exporter keeps the first unit registered for each metric name.
Platform and Collection-Mode Compatibility
Choose a collection mode below to compare Linux, macOS, and Windows. These
tables apply when the category is enabled. They cover both polling and
on-demand collect(). Polling emits marks. collect() returns a snapshot
without emitting a mark.
The tables use these terms:
- Yes: Relay can collect the field, though a reading can still fail.
- If available: Relay needs the named system interface, driver, or process group.
- No: The field could apply, but Relay does not collect it on that system.
- N/A: The field does not apply to that mode or operating system.
Both No and N/A give a null field or an empty event list. A field
marked Yes can also be null if a process exits, access is denied, or a
counter cannot be read.
A Linux control group (cgroup) sets limits for a group of processes. An exclusive cgroup contains only the process or tree that Relay measures. A Windows Job Object also groups processes so Relay can read group limits.
Linux CPU and memory limits include restrictions from parent cgroups up to the visible cgroup mount. These limits are upper bounds. Other processes that share a parent cgroup can reduce the resources available to the measured process or tree. Relay cannot read restrictions above the visible mount.
Global
Application Process
Process Tree
This is the CLI default. Process totals cover processes visible to Relay.
Relay might see fewer processes than the machine runs. Check
process_sampling for the number of visible and sampled processes.
The following table shows what global collection can measure:
Global CPU and disk totals add readable values for each counter. CPU use and system memory use system-wide readings. Linux pressure values come from the system pressure interface. GPU process records include only visible processes that the driver can identify. Global mode has no owned process group, so group limits and descendant counts do not apply. It reports system memory instead of per-process memory.
GPU Vendor Support in Every Collection Mode
Device records describe the whole GPU, even in application-process or process-tree mode. Process records cover visible processes in global mode, the selected process ID in application-process mode, or the selected process and its children in process-tree mode. A driver can provide memory without compute use, or compute use without memory. Choose a vendor to compare operating systems. These tables apply to every collection mode.
NVIDIA
AMD
Intel
Apple
Other
This table shows NVIDIA GPU measurements by operating system:
A backend is the driver interface that provides GPU readings. Each backend has different requirements:
- NVIDIA: NVML needs a working NVIDIA driver and library. Process compute use also needs the optional process-use API and an earlier reading in the same polling or on-demand series.
- Linux AMD and Intel: Process readings need readable DRM
fdinfofiles and driver counters. Device readings come from Linux system files (sysfs). AMD device readings usemem_info_vram_usedfor memory andgpu_busy_percentfor compute use. Intel device memory needsmem_info_vram_used. Relay does not collect Intel device compute use on Linux. - macOS: Device readings use driver
PerformanceStatisticsvalues when available. Apple process compute use comes from changes in AGX GPU time. These diagnostic values can change between macOS versions. Relay needs theioregutility and permission to read the device registry. Apple process data also needs exactly one detected Apple GPU that matches the selector. - Windows AMD and Intel: Relay has no collector for these GPUs.
Relay cannot collect Apple, AMD, or Intel process GPU memory on macOS. Apple
process compute use needs an earlier reading. Relay caps it at 100% when
parallel GPU work overlaps. A missing backend or field gives no record or a
null measurement.
GPU memory has different meanings across backends. Linux AMD and Intel drivers report used video memory. NVIDIA NVML reports used device memory. macOS drivers can report video or system memory. Relay converts these readings to KiB before applying your output unit, but they do not measure the same kind of memory.
Compute percentages also cover different periods of activity. AMD reports
whole-GPU busy time, NVIDIA reports GPU use, and macOS reports device use.
Linux AMD and Intel drivers and macOS omit a device record when neither memory
nor compute use is available. NVIDIA NVML can return a device record with both
measurements null.
Limits That Apply to Every Mode
These conditions apply in every collection mode:
- Filesystem capacity needs at least one configured
filesystem_pathsentry and a successful query for that path. An empty list produces an empty array. - Process I/O needs
disk.process_io = true. macOS currently reports bytes, but not read or write operation counts. - Device-wide GPU records need
gpu.device_metrics = true. Process GPU records needgpu.process_metrics = true. Both groups also need driver access. - GPU process utilization and CPU consumption rate need two readings in the
same polling or on-demand series. Disk and network throughput also need two
readings. Their first values are
null.
By default, CPU and pressure times use milliseconds, memory uses KiB, disk transfers and capacity use bytes, network data uses bytes, and GPU use is a percentage. Operating systems can define similar counters differently, so compare them with care.
On platforms other than Linux, macOS, and Windows, the plugin cannot create a collection target and activation fails.
Polling Marks and Scope Behavior
When the component and polling are enabled, each poll collects one snapshot and
emits a nemo.relay.resource_metrics mark. The CLI pauses process-tree polls
until it has a launched process to measure. The mark uses Relay’s reserved
nemo.relay.metric_measurements schema. An enabled Agent Trajectory
Observability Format (ATOF) subscriber writes the mark through the usual event
path. An enabled OpenTelemetry metrics pipeline exports available measurements
as metrics. Refer to
ATOF
and OpenTelemetry for exporter
configuration.
Each mark includes nemo.relay.resource.sample_count with value 1 and unit
events. This records the poll even when the selected categories have no
available measurements on the host.
The same selected-scope sample is copied under each active explicit Agent
scope. A custom scope does not receive a copy. If there are no active Agent
scopes, Relay emits one mark under the runtime root. Copies have the same
measurements and timestamp. They help each trace see the sample; they do not
represent separate per-agent resource use and must not be added together.
Agent scope start and end events only update the list of active scopes. They do not trigger an extra sample. Each copy goes through the target scope’s subscriber and sanitizer settings.
The CLI uses global scope by default. Embedded applications use the process tree
rooted at the current process ID by default. This can include child processes
that the host application started for other work. In the CLI, explicitly
selecting process_tree measures the launched agent process and its active
descendants. The runtime chooses the process ID; plugin configuration does not.
Collect One Snapshot
On-demand collection requires an active resource_metrics component, but it
does not require polling. Await the collection call to get a snapshot. The call
does not emit a polling mark. The first sample in a sampling series has no
consumption_rate, because that value needs two samples to calculate a change
over time.
After a Unix fork, the child cannot use the parent’s active plugin or its
collection target. Start a new executable in the child and activate the plugin
there before calling collect().
Collection runs OS and driver queries on a worker so they do not block your async runtime. Rust requires a Tokio runtime, and Python requires an active async event loop. Cancelling the wait does not stop a query already running. That query may finish and update the baseline used to calculate rates.
Each example activates the plugin, collects one snapshot, and closes the activation. Polling stays disabled by default. Choose your language:
Rust
Python
Node.js
Add these dependencies to your project’s Cargo.toml:
Save the following code as src/main.rs, then run cargo run:
Python returns the timestamp as a timezone-aware datetime. Rust returns a
UTC DateTime. Node.js returns a Promise. Await it to get a snapshot with an
ISO string timestamp. Node.js uses bigint for integer measurements and
process-sampling counts above JavaScript’s safe-integer range.
For plugin configuration files, precedence, and CLI editing, refer to Plugin Configuration Files.