OpenTelemetry
Use the opentelemetry section to configure independent OpenTelemetry Protocol
(OTLP) trace, log, and metric pipelines. OpenTelemetry support is always
included; no Cargo feature enables or disables it.
Relay classifies each sanitized event before OTLP signal fan-out:
Scope events maintain lineage for log correlation, but Relay does not export them as log records or metric measurements. Metric marks do not fall back to logs or traces when metric export is disabled. A valid mark with the reserved metric schema routes exclusively to metrics. An invalid reserved metric mark reaches no OTLP signal and produces a rate-limited operational diagnostic. Other marks retain the existing projection-specific trace behavior and can be exported as logs.
The metric schema is a NeMo Relay routing contract, not an ATOF-wide metric
semantic. ATOF remains at version 0.1, treats data_schema as opaque, and
defaults mark data_schema to null. A future specification discussion can
evaluate standardizing metric routing, timestamps, aggregation, and exemplars.
Trace Projections
Each endpoint selects one fixed semantic projection:
You can repeat a type or combine types. Each endpoint owns an independent exporter and can use a different endpoint.
The gen_ai projection keeps the
OpenTelemetry GenAI semantic-conventions v1.42-era snapshot
as its core compatibility baseline. Relay also implements selected newer
Development attributes. Refer to
GenAI Projection Attribute Support for
the exact current inventory and its newer registry baseline.
NeMo Relay uses the currently vendored OpenTelemetry Rust SDK (0.32). It
deterministically derives
compliant trace and span IDs from Relay lifecycle UUIDs, so endpoints that
receive the same event stream use the same identifiers and parentage. Different
endpoint types must therefore use independent OTLP destinations; configuring
them with the same endpoint and transport is rejected to prevent identifier
collisions at the receiver. Duplicate detection compares canonical destinations:
HTTP and HTTPS default ports are realized, repeated and trailing path slashes
are normalized, and standardized loopback hosts such as localhost, names
under .localhost, 127.0.0.0/8, and ::1 are equivalent. Relay does not use
DNS resolution for this comparison, and query strings remain significant.
Log lineage follows the same root selection as traces. When a local parent was
not observed by the exporters, the observed scope starts a trace using its own
UUID; a propagation-root UUID alone does not establish an imported parent.
Marks under that scope, including late marks within the completed-context retention
window, use its trace identity. An explicit imported parent continues the propagated
trace. Log-to-span joins require both exporters to observe the corresponding
scope lifecycle.
Relay propagation continues the Relay-derived trace across the import boundary
by default. Use capture_rootless_propagation_context() only when a receiver
must start a new OpenTelemetry trace. Carry W3C traceparent and
tracestate alongside Relay propagation when an integration also needs to
preserve upstream OpenTelemetry sampling or vendor state.
plugins.toml Example
The following version-4 configuration exports a gen_ai trace and derives log
and metric destinations from the trace destination.
Relay replaces the terminal /v1/traces with /v1/logs and /v1/metrics.
Derived destinations copy the trace endpoint’s transport, authentication,
resource attributes, resource-metadata promotion prefixes, service identity,
instrumentation scope, and timeout. In this example, all three signals use the
same authorization and nv.project routing value.
When enabled = true, configure at least one trace endpoint or an enabled
signal with explicit endpoints. NeMo Relay constructs every endpoint before
registering subscribers. An invalid endpoint is skipped with an activation
warning while valid endpoints register. Activation fails when no trace, log,
or metric endpoint can be registered. A delivery failure from one exporter does
not stop application work or delivery to the other exporters.
Trace Endpoint Fields
All trace projections (full, gen_ai, and openinference) require HTTPS for
remote collectors, with either HTTP or gRPC transport. HTTP is permitted only
for localhost and loopback IP addresses such as 127.0.0.1 and [::1].
Redaction does not disable this requirement. Existing remote plaintext trace
endpoints must migrate to HTTPS or use a loopback collector.
OTLP/HTTP trace exporters do not follow redirects, even to another HTTPS endpoint. This keeps an export from bypassing transport checks or sending its payload and credentials to another destination. Use the final collector URL. Log and metric HTTP exporters do the same when any header source is configured. gRPC does not use HTTP redirects.
Event Metadata Promotion
Wrapped CLI events can include trusted agent_version metadata. To export this
generic field as a span attribute, include "agent_version" in
promote_metadata_prefixes. Refer to
Coding-Agent Identity Metadata
for its production and lifecycle semantics.
Set promote_metadata_prefixes on a trace endpoint to copy selected keys from
the final sanitized Event metadata into that endpoint’s OpenTelemetry output.
The setting defaults to an empty list, so Relay does not promote metadata unless
you configure at least one prefix.
For example, "app.", "app_", and "app-" are valid literal prefixes.
Leading or repeated dots, whitespace, other punctuation, and glob expressions
such as "app.*" are rejected.
Matching is case-sensitive and compares the beginning of each key literally.
Relay does not infer a dot or metadata-key segment boundary. For example,
"app." selects app.name and app.version, but not app_name. The broader
"app" prefix selects all three keys.
Configure the narrowest prefix that selects the metadata you intend to export.
Scope-start and Scope-end are separate Event records. When Scope-end completes
the span, a metadata key present on that Event replaces the corresponding
promoted Scope-start value. Mark metadata is promoted to the attributes of the
projected span event or tool span. The gen_ai projection continues to omit
Marks.
Promotion supports strings, booleans, signed 64-bit integers, floating-point
numbers, empty arrays, and homogeneous arrays of those primitive types. Relay
omits rejected values and records one bounded runtime diagnostic per rejected
key. The diagnostic code is
otel.metadata_promotion_value_unsupported.<metadata-key>, its message contains
the key and rejection reason, and its count is the number of occurrences for
that key. Relay does not record the rejected value or stop trace export. Match
the otel.metadata_promotion_value_unsupported. prefix to monitor rejected span
metadata keys.
Projection-owned attributes take precedence over promoted metadata with the
same key. For full and openinference, configured attribute-mapping aliases
also take precedence. Relay also omits selected
keys in namespaces owned by Relay or supported semantic projections:
nemo_relay., gen_ai., error., exception., input., output., llm.,
openinference., server., service., session., tool., tool_call., and
user.. Relay omits the bare metadata key as well. Rejected keys produce a
rate-limited operational diagnostic without dropping the Event or span.
Promotion does not modify the Event or ATOF payload. In OTLP trace output only,
Relay removes successfully promoted keys from serialized Relay metadata
attributes. This applies to every Scope-start, Scope-end, and Mark event, and
also to OpenInference’s metadata JSON attribute. Resource-promotion prefixes
participate in that filtering for every event even though only a trace root’s
Scope-start metadata can create resource attributes. Keys that cannot be
promoted, including values overridden by configured resource attributes, remain
in serialized metadata. Use resource_attributes instead for static values
that must be attached to every span from an endpoint.
Root Resource Metadata Promotion
Set promote_resource_metadata_prefixes on a trace, log, or metric endpoint to
derive resource attributes from the sanitized metadata on a root Scope-start
Event. Every span, scoped log, and scoped metric measurement in that scope tree uses
the same resource for its configured endpoint; child metadata and later root
metadata changes do not modify it. Derived log and metric endpoints inherit the
trace endpoint’s prefixes. The prefixes and supported value types match
promote_metadata_prefixes.
For explicit log and metric endpoints, the Python plugin helper
OpenTelemetrySignalEndpointConfig and the Node.js plugin helper
openTelemetrySignalEndpoint accept promote_resource_metadata_prefixes.
The Go plugin helper ObservabilityOpenTelemetrySignalEndpointConfig exposes
PromoteResourceMetadataPrefixes. The Go trace endpoint helper
ObservabilityOpenTelemetryEndpointConfig exposes the same field, which derived
log and metric endpoints inherit. Omit the setting to keep promotion disabled.
These plugin settings also apply when endpoints are configured through TOML or
JSON. The Python, Node.js, and Go direct log and metric subscriber APIs do not
expose this option.
Configured service_name, service_namespace, service_version, and explicit
resource_attributes take precedence over promoted metadata with the same key.
Each distinct effective resource creates a retained OTLP provider pipeline for
each enabled signal, so use only controlled, low-cardinality values such as
deployment, region, client version, or environment identity. Do not promote
request, tenant, or user identifiers.
Each log or metric endpoint retains at most 16 dynamic resource pipelines,
plus its configured base provider. Existing resource keys continue to reuse
their pipelines at the limit; pipelines are not evicted. Additional resource
keys use the base resource and produce otel.resource_metadata_pipeline_limit.
This preserves delivery but omits the promoted resource distinction, so metric
measurements that differ only by that resource can be aggregated together.
The limit is fixed, applies independently to each log and metric endpoint, and
does not cap trace pipelines or total process memory.
Each log or metric endpoint tracks at most 4,096 active resource scopes,
including nested scopes. Once full, it keeps existing routes and declines new
scope starts with otel.resource_metadata_active_scope_limit; marks attached
to those untracked scopes use the base resource. Ending a tracked scope frees a
slot for a subsequent scope start. Active routes are not evicted by age, so
long-running scopes retain their identity. Missing end events can occupy slots
until the subscriber is replaced, but cannot grow this active-route map beyond
the limit. This limit does not bound completed-route caches, trace correlation
state, or total process memory.
Completed resource routes expire before processing an event beyond the TTL;
events exactly at the boundary remain linked. Logs use
opentelemetry.logs.completed_span_context_ttl_millis (default: 60 seconds).
Metrics use a fixed 60-second resource-route TTL. Expiration is driven by event
timestamps and runs when another event arrives, rather than on a background
timer.
Relay records bounded runtime diagnostics for this setting. A rejected value
produces otel.resource_metadata_promotion_value_unsupported.<metadata-key>.
A failed resource-pipeline construction produces
otel.resource_metadata_pipeline_build_failed; Relay then exports the affected
signal through the endpoint’s configured resource instead. Match both the
otel.metadata_promotion_value_unsupported. and
otel.resource_metadata_promotion_value_unsupported. prefixes to monitor all
rejected metadata keys.
Log and Metric Endpoint Resolution
An enabled logs or metrics section can omit endpoints. Relay then derives
one signal endpoint from every trace endpoint:
- A bare HTTP authority, with or without a root trailing
/, gains/v1/logsor/v1/metrics. - A terminal
/v1/traces, including one below a path prefix, is replaced with the signal path. Query parameters are preserved. - A gRPC endpoint reuses its authority without path rewriting.
- A trace endpoint with any other custom path cannot be derived. Configure signal endpoints explicitly in that case.
An explicit nonempty signal endpoint list replaces derivation. Relay preserves
an explicit custom signal path exactly, but rejects an obvious standard path
for another signal, such as /v1/traces in a log endpoint. An explicit empty
list is invalid when the signal is enabled.
The following example sends logs to a custom intake path while metrics continue to derive from the trace endpoint:
The signal endpoint fields are endpoint, transport, headers,
header_env, header_file, resource_attributes,
promote_resource_metadata_prefixes, service_name, service_namespace,
service_version, instrumentation_scope, and timeout_millis. Their defaults
match the corresponding trace fields. Each signal rejects duplicate
destinations within that signal. Logs, metrics, and traces can share the same
authority because OTLP treats them as different signals. All three signals
automatically include the reserved telemetry.sdk.name, telemetry.sdk.language,
and telemetry.sdk.version resource attributes; configuring those keys in
resource_attributes rejects the endpoint.
Trace Batch Processor Configuration
Configure batch processing independently on each endpoint with
max_queue_size, max_export_batch_size, and scheduled_delay_millis.
When an endpoint omits a field, the corresponding standard OpenTelemetry
environment variable applies process-wide. If neither is set, the SDK default
applies.
The precedence for each setting is endpoint value, then environment variable, then SDK default. Set environment variables before the plugin activates.
Endpoint values must be positive integers. If both endpoint size fields are
set, max_export_batch_size must not exceed max_queue_size. When one size is
inherited, the SDK caps the effective batch size at the effective queue size.
The SDK falls back to its default for malformed environment values. Queue and
batch sizes count spans, not bytes.
Completed Scope Lineage Retention
Trace endpoints retain a completed scope’s trace and parent span context for
completed_span_context_ttl_millis after its scope-end event. A late mark in
that window remains attached to the original trace and parent span. When the
TTL expires, Relay emits subsequent marks as orphan spans and records the
otel.completed_span_context_expired runtime diagnostic when it purges expired
contexts.
Keep closed scope handles only for short deferred follow-up work. Prefer emitting an event before the scope closes, or create a new active scope for later work. Increasing the TTL retains more completed contexts: memory grows with the completed-scope rate, TTL, and number of configured trace endpoints.
OTLP logs use the same TTL-based completed-scope lineage behavior. Configure
opentelemetry.logs.completed_span_context_ttl_millis independently when log
export is enabled. Metrics retain completed resource routes for 60 seconds when
resource promotion is enabled, but do not retain trace-correlation context. Do not
rely on a closed scope handle for long-running follow-up work; emit the mark
before closing the scope or use an active/new scope instead.
If an endpoint’s explicit batch settings are invalid, Relay skips that endpoint
and records an observability.invalid_otel_endpoint configuration warning with
its opentelemetry.endpoints[N] field. Other valid endpoints continue to
activate. Activation still fails when no trace, log, or metric endpoint can be
registered.
Relay also skips and logs any trace endpoint that fails during exporter construction. This includes malformed collector destinations that cannot be detected during configuration validation.
Relay’s thread-based batch processor exports serially, so it does not expose
the SDK’s concurrent-export setting. It also does not expose a separate batch
processor export timeout; use the endpoint’s timeout_millis to bound each
OTLP request.
Known limitation: OTLP partial success is not reported. A collector can
return a successful OTLP response while rejecting individual spans, log
records, or metric data points. With the vendored OpenTelemetry exporter,
Relay treats that response as successful: it does not add a runtime diagnostic,
and force_flush() and shutdown() can succeed. Monitor collector-side logs
and rejection metrics when investigating missing telemetry.
A full queue drops completed spans instead of applying backpressure to application work. Bursts can therefore drop spans that finish late, including an enclosing root span, and leave an incomplete trace in the backend.
The OpenTelemetry SDK logs
BatchSpanProcessor.SpanDroppingStarted at warning level when each endpoint
first drops a span. It suppresses additional first-drop warnings for that
endpoint to avoid a log storm. During graceful shutdown, it logs
BatchSpanProcessor.SpansDropped with the endpoint processor’s exact
dropped_span_count and max_queue_size.
For plugin-managed exporters, NeMo Relay also records otel.spans_dropped in the
active plugin report’s runtime_diagnostics. Its count is the exact number
of dropped spans, field identifies the affected
opentelemetry.traces[N].endpoint, and message includes the configured
endpoint origin (scheme, host, and port), without URL credentials, paths, query
parameters, or fragments. If spans were dropped, clearing the plugin returns a delivery
failure error and retains the diagnostic for inspection. This error does not
disable later plugin configuration.
Increasing an endpoint’s max_queue_size, or the process-wide
OTEL_BSP_MAX_QUEUE_SIZE fallback, can reduce the risk for a known burst size,
but a finite queue does not guarantee lossless telemetry. Always clear the
plugin during graceful shutdown so NeMo Relay can record the final drop count
and the SDK can attempt to export queued spans.
Endpoint Capacity and Sizing
NeMo Relay does not impose a maximum number of OpenTelemetry endpoints. Each trace, log, or metric endpoint owns an exporter, signal provider, processor or reader, and exporter runtime resources. Nothing in those export stacks is shared between endpoints. Queue capacity and memory are per endpoint, and total export traffic grows with the endpoint count.
Typical deployments need one to three endpoints. Validate configurations with tens or hundreds of endpoints against the process limits for threads, memory, and network egress before deploying them.
Use header_env for secrets so configuration files contain only environment
variable names. Each variable contains the complete header value. NeMo Relay
validates variable names without reading their values, then resolves and
snapshots the values when the plugin activates. Every referenced variable name
must be nonblank and have no surrounding whitespace. Its value must be set and
nonblank, with no surrounding whitespace, when the component activates. A header name
cannot appear in both headers and header_env, including names that differ
only by ASCII case. Reactivate the plugin to pick up a changed environment
value.
Use header_file when another process updates a credential in place, such as a
projected token file. Each file must be regular and exist when the plugin
activates. Relay reads the file only when it exports. Relay removes trailing
whitespace. A missing, unreadable, blank, or invalid value fails only that
export, and diagnostics do not show the value. A header can use only one of
headers, header_env, and header_file. Header names are case-insensitive.
This applies to trace, log, and metric endpoints.
Remote trace endpoints must use HTTPS. When you set any header source
(headers, header_env, or header_file), remote log and metric endpoints
must also use HTTPS. HTTP is allowed only for localhost and loopback IP
addresses. OTLP/HTTP exporters do not follow redirects when headers are set.
For example, an OIDC token-file writer must write the complete value Bearer eyJ... to the file. Relay does not add the authentication
scheme for you.
Process-global OTEL_EXPORTER_OTLP_HEADERS,
OTEL_EXPORTER_OTLP_TRACES_HEADERS, OTEL_EXPORTER_OTLP_LOGS_HEADERS, and
OTEL_EXPORTER_OTLP_METRICS_HEADERS are rejected because they cannot be
isolated between endpoints. Put non-secret values in each endpoint’s headers
map and secret variable references in header_env.
full and openinference endpoints retain the legacy mark and attribute-alias
controls shown above. gen_ai is standards-only: it ignores those controls and
does not emit Relay-private attributes. semantic_selector and
capture_content are unsupported.
On a successful tool end span with a present, non-null annotation, the full
and openinference projections emit the opaque value as one JSON string
attribute named nemo_relay.tool.result.annotation. Relay does not flatten the
annotation’s application-defined keys. The gen_ai projection omits this
Relay-private attribute.
Emit Log and Metric Marks
Rust, Python, and Node.js generic mark APIs accept optional data_schema and
severity values. Prefer the typed metric helper rather than constructing the
reserved schema by hand: metric in Rust, Python, and Node.js. The helper
validates the complete measurement group before publishing the mark. The
exporter validates the sanitized payload again before recording any measurement.
The following examples emit one warning log mark and one metric mark:
Python
Node.js
Rust
Mark sanitizers run for both calls. Routing uses the immutable data_schema
after sanitization, and a metric mark never falls back to the log pipeline.
Log Export
The log pipeline exports one OTLP LogRecord for each sanitized non-metric
mark. Marks with data_schema = null and marks with an application-defined
schema are logs. Any mark that uses the reserved
nemo.relay.metric_measurements schema name is routed away from logs,
including unsupported schema versions and invalid metric payloads. Scope start
and end events update the lineage used for correlation but do not become log
records.
Use the typed severity argument on the generic mark API. Relay stores it in
sanitizer-visible metadata as nemo_relay.log.severity. The typed argument
overrides that metadata key and requires metadata to be an object. After mark
sanitizers run, Relay parses the remaining key, defaults an absent key to
info, and drops a log record with an invalid value. Supported values are
trace, debug, info, warn, and error; warning is accepted as an
alias for warn.
The logs section applies these processing settings to every log endpoint:
Relay maps the event timestamp to the log timestamp and post-sanitization
processing time to the observed timestamp. Sanitized data becomes the
structured body. An absent or top-level JSON null payload has no body; a
nested JSON null becomes the string "null" because OTLP AnyValue has no
null variant.
The log attributes preserve the mark name, UUID, optional parent UUID,
category and category profile, schema, sanitized metadata, and ATOF version
under nemo_relay.* keys. A mark in a resolvable active or completed scope
receives trace and span context. An orphan mark receives no invented trace
context. Relay leaves OTLP event_name unset with the currently vendored
OpenTelemetry SDK (0.32) and retains the dynamic name in
nemo_relay.mark.name.
Telemetry logs are separate from NeMo Relay’s operational stderr and file
logging. minimum_severity does not inherit NEMO_RELAY_LOG, and operational
diagnostics are not fed back into ATOF or OTLP.
Metric Export
The metric pipeline consumes only sanitized marks with this exact schema:
A mark with the reserved schema name and an unsupported version or invalid payload is dropped from both logs and metrics. Relay emits a rate-limited operational diagnostic without creating another ATOF or OTLP event.
The measurements array is required and nonempty, and unknown fields are
rejected. Each measurement is an SDK recording operation, not a pre-aggregated
OTLP point:
Unsigned values must not exceed i64::MAX, which prevents loss in the pinned
OTLP conversion. Metric names must be 1 to 255 ASCII bytes, start with a letter,
and contain only letters, digits, _, ., -, or /. Units must be ASCII
and at most 63 bytes. Optional histogram boundaries can include negative
values, but every boundary must be finite, strictly increasing, unique, and the
list can contain at most 64 entries.
Attributes can contain strings, Booleans, signed integers, finite doubles, and
homogeneous nonempty arrays of those primitive types. Blank keys, nulls,
nested objects, mixed arrays, and unsigned integers above i64::MAX are
invalid. Do not use event UUIDs, timestamps, metadata, trace IDs, or other
high-cardinality values as metric attributes.
Relay treats the complete mark atomically. It records no measurements from the mark when a measurement is invalid, an instrument descriptor conflicts, or an instrument limit would be exceeded. Instrument names compare case-insensitively, and a name must retain its kind, numeric type, unit, description, and histogram boundaries for the lifetime of that destination.
The metrics section applies these settings to every metric endpoint:
Metric points use SDK collection timestamps. The currently vendored
OpenTelemetry SDK (0.32) cannot preserve the source mark timestamp or attach
a trace-linked exemplar through this path. Relay does not emulate correlation
with high-cardinality attributes.
GenAI Projection
Set an OpenTelemetry endpoint’s type to gen_ai to select this projection:
The gen_ai endpoint uses these operation names:
Marks are omitted. Relay scope types without GenAI semantics are emitted as
minimal internal spans so that the original span parentage is preserved. This
projection never emits nemo_relay.* fields. LLM spans include the
gen_ai.system_instructions, gen_ai.input.messages, and
gen_ai.output.messages attributes as JSON strings that follow the
OpenTelemetry GenAI schemas. Each attribute is emitted only when its normalized
instructions, messages, or response content is present. Redact sensitive
content with an LLM or event sanitizer. Tool content is exported by default;
retrieval payloads are not exported. Set top-level enable_full_payloads = true to retain complete
sanitized LLM request history on every start span.
GenAI Projection Attribute Support
This inventory follows the 72-attribute OpenTelemetry GenAI registry snapshot used for the current Relay audit. Every listed upstream attribute is marked Development. This newer inventory includes attributes added or renamed after Relay’s v1.42-era compatibility baseline.
Supported means Relay’s native GenAI projection emits the exact attribute on
an applicable signal when an authoritative normalized or instrumentation source
exists. It does not mean every integration supplies that optional source.
Partial means Relay emits the attribute but cannot represent the complete
upstream cardinality or value-shape contract. Not supported means the native
projection does not emit it. The gen_ai.* namespace is reserved, so generic
metadata promotion cannot be used to add an unsupported GenAI attribute.
| Projection Attribute | Support | Description and Notes |
|---|---|---|
gen_ai.agent.description | Supported | Application-provided free-form agent description. Relay emits it on Agent scopes and marked CLI turns when canonical or recognized alias metadata supplies it. |
gen_ai.agent.id | Not supported | Stable identifier of a hosted GenAI agent resource. Relay does not yet distinguish hosted-agent create or client operations, and it does not substitute transient scope, session, subagent, or harness identifiers. |
gen_ai.agent.name | Supported | Application-provided human-readable agent name. Relay uses Agent scope identity and explicit applicable turn or tool metadata; marked turns omit it when no authoritative name exists. |
gen_ai.agent.version | Not supported | Version of a hosted GenAI agent. Relay does not yet classify hosted-agent operations. The generic agent_version captured by wrapped CLI launches identifies the harness executable and is intentionally not projected as this attribute. |
gen_ai.conversation.compacted | Not supported | Indicates that the effective context is a compacted view of an earlier conversation. Relay does not yet propagate positive compaction state into the later LLM operation; false must not be emitted. |
gen_ai.conversation.id | Supported | Stable conversation, session, or thread identifier. Relay projects the canonical key or recognized conversation, session, and thread aliases on applicable Agent, turn, LLM, and tool spans. |
gen_ai.data_source.id | Supported | Identifier of the GenAI data source. Relay projects it on retriever spans from the canonical key or recognized data-source aliases; prefer the GenAI system identifier over an external storage name. |
gen_ai.embeddings.dimension.count | Supported | Requested output embedding dimension count. Relay emits a positive integer on embedder spans from the canonical key or dimensions. |
gen_ai.evaluation.explanation | Not supported | Evaluator-provided explanation for an assigned score. The convention places it on a gen_ai.evaluation.result event; Relay has no corresponding standard event projection. |
gen_ai.evaluation.name | Not supported | Name of the evaluation metric. Relay has no gen_ai.evaluation.result event projection. |
gen_ai.evaluation.score.label | Not supported | Human-readable interpretation of an evaluation score. Relay has no gen_ai.evaluation.result event projection. |
gen_ai.evaluation.score.value | Not supported | Numeric evaluation score. Relay has no gen_ai.evaluation.result event projection. |
gen_ai.input.messages | Supported | Chat history supplied to the model. Relay serializes retained normalized and sanitized messages as schema-shaped JSON text. enable_full_payloads controls complete request-history retention, not whether present content can be projected. |
gen_ai.memory.query.text | Not supported | Search query used to retrieve memories. Relay has no standard memory-operation lifecycle or sensitive-content opt-in contract for this value. |
gen_ai.memory.record.count | Not supported | Number of memory records relevant to an operation. Relay has no normalized memory-result lifecycle. |
gen_ai.memory.record.id | Not supported | Unique memory-record identifier. Relay has no normalized memory-record identity contract. |
gen_ai.memory.records | Not supported | Memory records stored or retrieved by a memory operation. Relay has no memory lifecycle, normalized record shape, or sensitive-content opt-in contract. |
gen_ai.memory.store.id | Not supported | Unique identifier of a memory store. Relay has no standard memory scope or operation contract. |
gen_ai.operation.name | Supported | Name of the GenAI operation. Relay projects invoke_agent, chat, generate_content, text_completion, execute_tool, embeddings, or retrieval on the corresponding scopes and marked CLI turn roots. |
gen_ai.output.messages | Partial | Model output messages, with one message per returned choice or candidate. Relay currently normalizes and projects only one assistant candidate, so additional parallel generations are not retained. |
gen_ai.output.type | Supported | Output modality requested by the client. Exact canonical metadata wins; Relay otherwise derives json, text, or speech from supported OpenAI request shapes and omits ambiguous, multiple, or unsupported modalities. |
gen_ai.prompt.name | Not supported | Name that uniquely identifies a prompt template. Relay has no authoritative cross-provider prompt identity source. |
gen_ai.prompt.variable.<name> | Not supported | Runtime value supplied for a named prompt-template variable. Relay has no normalized variable map or explicit sensitive-content policy for these dynamic attributes. |
gen_ai.prompt.version | Not supported | Version of the prompt template. Relay has no authoritative cross-provider prompt-version source. |
gen_ai.provider.name | Supported | GenAI provider identified by the instrumentation. Relay uses explicit canonical or alias metadata, recognized routes and event names, or the normalized provider API; it can emit custom values such as oci.genai. |
gen_ai.request.choice.count | Supported | Requested number of candidate completions. Relay projects non-default OpenAI Chat n values; the default value of one is omitted. |
gen_ai.request.encoding_formats | Supported | Requested embedding encoding formats. Relay projects the canonical key or recognized singular and plural aliases on embedder spans. |
gen_ai.request.frequency_penalty | Supported | Request frequency-penalty setting. Relay projects it from normalized OpenAI Chat requests. |
gen_ai.request.max_tokens | Supported | Maximum number of tokens requested for generation. Relay projects the normalized request value. |
gen_ai.request.model | Supported | Name of the model targeted by the request. Relay projects normalized request or authoritative model metadata on applicable spans. |
gen_ai.request.presence_penalty | Supported | Request presence-penalty setting. Relay projects it from normalized OpenAI Chat requests. |
gen_ai.request.previous_response.id | Supported | Identifier of a prior response used as context for the current operation. Relay projects the normalized previous-response ID. This attribute is newer than the v1.42-era baseline. |
gen_ai.request.reasoning.level | Supported | Requested reasoning or thinking effort. Relay uses exact canonical metadata, OpenAI Chat reasoning_effort, or normalized reasoning effort. |
gen_ai.request.seed | Supported | Seed intended to make repeated requests more deterministic. Relay projects it from normalized OpenAI Chat requests. |
gen_ai.request.stop_sequences | Supported | Sequences that stop further token generation. Relay projects the normalized request list. |
gen_ai.request.stream | Supported | Indicates a streaming request. Relay emits only true; absence represents a non-streaming request as required by the convention. |
gen_ai.request.stream_cursor | Not supported | Cursor used to resume a streamed response after the last received event. Relay has no fetch or resume operation classification or normalized cursor source. |
gen_ai.request.temperature | Supported | Request temperature setting. Relay projects the normalized request value. |
gen_ai.request.top_k | Supported | Top-K sampling limit used during generation. Relay projects normalized Anthropic top_k; it does not misclassify OpenAI top_logprobs as this attribute. |
gen_ai.request.top_p | Supported | Request nucleus-sampling setting. Relay projects the normalized request value. |
gen_ai.response.finish_reasons | Partial | Ordered reasons that each returned generation stopped. Relay retains and emits one normalized finish reason, so it cannot represent multiple candidates or an expected generation that ended before producing a normal reason. |
gen_ai.response.id | Supported | Unique completion or response identifier. Relay projects the normalized response ID. |
gen_ai.response.model | Supported | Name of the model that generated the response. Relay projects normalized LLM response data or an authoritative embedder response source. |
gen_ai.response.status | Not supported | Provider-reported lifecycle status when a response is fetched or polled. Relay does not classify fetch or poll operations and intentionally does not copy ordinary inference status into this attribute. |
gen_ai.response.time_to_first_chunk | Supported | Seconds from managed request execution to the first received provider protocol chunk. This is not time to first text token. Relay omits it for non-streaming calls and streams with no chunk, and emits the companion standard histogram only when provider identity is known. |
gen_ai.retrieval.documents | Not supported | Documents returned by retrieval. Relay has no normalized document-result source or content policy; raw retrieval payloads are not exported by the GenAI projection. |
gen_ai.retrieval.query.text | Not supported | Query text used for retrieval. Relay has no explicit normalized query source or sensitive-content opt-in contract. |
gen_ai.retrieval.top_k | Supported | Maximum number of requested retrieval documents. Relay projects the canonical key or top_k on retriever spans. |
gen_ai.system_instructions | Supported | System instructions supplied separately from chat history. Relay serializes retained normalized and sanitized instructions as schema-shaped JSON text when present. |
gen_ai.token.type | Not supported | Token category used by the standard gen_ai.client.token.usage metric. It is not a span attribute, and Relay does not automatically emit that metric; custom typed metric marks can carry it independently. |
gen_ai.tool.call.arguments | Partial | Parameters passed to a tool call. Relay exports every non-null sanitized value as canonical JSON text. OpenTelemetry expects object-shaped data, so arrays and scalar values are preserved but do not satisfy the upstream schema. |
gen_ai.tool.call.id | Supported | Tool-call identifier. Relay prefers the typed call ID, then explicit instrumentation metadata, and never derives identity from argument or result payloads. |
gen_ai.tool.call.result | Partial | Successful result returned by a tool call. Relay omits failed results and exports every non-null sanitized value as canonical JSON text. OpenTelemetry expects object-shaped data, so arrays and scalar values do not satisfy the upstream schema. |
gen_ai.tool.definitions | Supported | Tool definitions available to the agent or model. Relay projects minimal schema-shaped definitions containing required type and name identities while omitting optional descriptions and parameter schemas. |
gen_ai.tool.description | Supported | Tool description. Relay emits nonblank explicit instrumentation metadata and never infers it from invocation arguments. |
gen_ai.tool.name | Supported | Name of the tool used by the agent. Relay uses explicit tool identity metadata or the tool scope name. |
gen_ai.tool.type | Supported | Type of tool used by the agent, such as function, extension, or datastore. Relay uses explicit metadata or source-backed harness defaults and preserves explicit refinements. |
gen_ai.usage.audio.cache_read.input_tokens | Not supported | Audio input tokens served from a provider cache. Relay’s normalized usage contract has no audio and cache modality breakdown. |
gen_ai.usage.audio.input_tokens | Not supported | Audio input-token count. Relay’s normalized usage contract has no audio modality breakdown. |
gen_ai.usage.audio.output_tokens | Not supported | Audio output-token count. Relay’s normalized usage contract has no audio modality breakdown. |
gen_ai.usage.cache_read.input_tokens | Supported | Input tokens served from a provider-managed cache. Relay projects normalized cache-read usage and includes known provider cache accounting in the aggregate input count. |
gen_ai.usage.cache_write.input_tokens | Supported | Input tokens written to a provider-managed cache. Relay projects normalized cache-write usage. This newer spelling replaces gen_ai.usage.cache_creation.input_tokens from the v1.42-era registry snapshot. |
gen_ai.usage.image.cache_read.input_tokens | Not supported | Image input tokens served from a provider cache. Relay’s normalized usage contract has no image and cache modality breakdown. |
gen_ai.usage.image.input_tokens | Not supported | Image input-token count. Relay’s normalized usage contract has no image modality breakdown. |
gen_ai.usage.image.output_tokens | Not supported | Image output-token count. Relay’s normalized usage contract has no image modality breakdown. |
gen_ai.usage.input_tokens | Supported | Total GenAI input-token count. Relay uses normalized usage and includes separately reported provider cache reads and writes when required to form the total. |
gen_ai.usage.output_tokens | Supported | Total GenAI output-token count. Relay projects normalized completion usage. |
gen_ai.usage.reasoning.output_tokens | Supported | Output tokens used for reasoning or extended thinking. Relay uses exact canonical metadata, OpenAI Responses reasoning details, or Gemini thought-token usage. |
gen_ai.usage.text.cache_read.input_tokens | Not supported | Text input tokens served from a provider cache. Relay’s normalized usage contract has no text and cache modality breakdown. |
gen_ai.usage.text.input_tokens | Not supported | Text input-token count. Relay’s normalized usage contract has no text modality breakdown. |
gen_ai.usage.text.output_tokens | Not supported | Text output-token count. Relay’s normalized usage contract has no text modality breakdown. |
gen_ai.workflow.name | Not supported | Application-provided low-cardinality workflow name. Relay has no Workflow scope or invoke_workflow operation classification. |
Usage is recorded per model request. Relay’s normalized response usage retains
an optional uncached_input_tokens count only when the provider’s response
contract proves it. The standard GenAI export deliberately does not add a
nonstandard uncached field: consumers can use the total
gen_ai.usage.input_tokens with the separately reported cache-read and
cache-write counts. An omitted cache or uncached count means unavailable, not
zero.
For managed streaming LLM calls, the end span also includes
gen_ai.response.time_to_first_chunk in seconds. Relay measures from the
start of managed stream execution until the first provider protocol chunk is
received; this is not a first-text-token measurement. The attribute is omitted
for non-streaming calls and streams that never yield a chunk.
The metric endpoint records the same sample as the standard
gen_ai.client.operation.time_to_first_chunk f64 histogram with unit s
when Relay can determine the provider dimension.
The gen_ai projection includes sanitized gen_ai.tool.call.arguments at tool
start, gen_ai.tool.call.result on successful completion, and
gen_ai.tool.definitions on inference spans by default. These attributes can
contain sensitive information. Use the PII redaction trajectory_context
preset to remove opaque payloads while retaining analytical structure and trace
parentage. Credential removal and event sanitizers run before projection.
enable_full_payloads controls LLM request-history retention independently.
Arguments and results are emitted as canonical JSON strings. Relay parses
serialized JSON before projection, then preserves every non-null sanitized JSON
value, including objects, arrays, strings, booleans, and numbers. Failed tool
calls do not emit a result attribute. Result annotations remain separate and
are never included in this projection. Tool definitions include only required
type and name properties. Known OpenAI Responses built-ins use their native
type as identity; recognized Gemini native tool groups emit an identity for each
tool in the group. Explicit native names are preserved. Unknown unnamed native
definitions are omitted rather than assigned invented identities.
Tool Identity
Tool identity comes from instrumentation metadata and the typed tool-call ID,
never from similarly named fields in tool arguments or results. Nonblank string
metadata supports gen_ai.tool.name, gen_ai.tool.type / tool_type,
gen_ai.tool.call.id / tool_call_id, gen_ai.tool.description /
tool_description / description, and gen_ai.agent.name / agent_name.
The typed call ID takes precedence. Absent tool names use the scope name;
unknown descriptions, types, and agent names are omitted.
Claude Code, Codex, and Pi tool hooks default gen_ai.tool.type to function
because these harnesses execute tools locally. They default gen_ai.agent.name
to claude-code, codex, or pi, or subagent:<id> when ownership is known.
These defaults apply to paired hooks and post-only hooks that synthesize a
start, without replacing explicitly supplied nonblank string metadata, including
the tool_type and agent_name aliases. Generic gateway sessions
do not infer a tool type or root executing-agent name. Invocation arguments
are never used to infer tool descriptions.
Harness defaults carry internal provenance. Explicit completion metadata may refine inferred tool type and agent name, including through their aliases. Explicit start metadata and typed tool-call IDs remain authoritative; later inferred values cannot replace them. This provenance is not exported as a GenAI attribute.
Claude Code and Codex MCP tool hooks with a qualified
mcp__<server>__<tool> name receive mcp.method.name = "tools/call" in event
metadata, including post-only hooks. Relay preserves existing method metadata
and does not infer MCP identity for ordinary tools, Pi, or generic gateway events.
The original qualified tool name remains unchanged. To include the method in
OTLP span attributes, configure promote_metadata_prefixes to include
"mcp.method.name".
This identifies a requested MCP tool operation, not proof that a request reached
the server or succeeded. Permission-denied calls retain their denial and error
metadata. The server segment is a configured alias; Relay does not derive
server.address, mcp.session.id, or connection status from it. Hook duration
measures the harness tool scope, not necessarily MCP transport latency.
Error Type Mapping
For managed LLM, tool, and stream failures, NeMo Relay maps structured
FlowError values to the OpenTelemetry error.type attribute:
External application and callback exceptions that do not have a more specific
FlowError classification emit internal_error. Python and JavaScript callback
boundaries also preserve the exception class separately, and both the full
and gen_ai projections emit an exception span event with exception.type.
NeMo Relay does not inspect error messages to recover exception class names.
When an errored parent span has no useful classification of its own, it
inherits the failed descendant’s error.type and exception type. When no
structured FlowError is available, such as a cancellation or dropped
execution, the projection emits _OTHER. Caller-provided error.type and
exception.type metadata take precedence over values derived from FlowError.
An authenticated Claude Code or Codex permission rejection closes the matched
active tool span with error.type = "guardrail_rejected". The corresponding
permission guardrail events carry the canonical gen_ai.tool.call.id in event
metadata so subscribers can correlate the decision with that tool call.
FlowError is an exhaustive Rust enum. Rust callers upgrading to this release
must handle the new CallbackException variant in exhaustive matches. It maps
to the same internal status as Internal, while retaining exception_type for
observability projection.
Direct Subscribers
Python
Node.js
Rust
Set each referenced environment variable before constructing the subscriber.
Direct trace, log, and metric configs resolve header_env when the subscriber
is constructed and retain that value for the subscriber’s activation. Changing
the process environment affects only a subsequently constructed subscriber.
Static headers remain unchanged. A header name cannot appear in both maps,
including names that differ only by ASCII case, and names within header_env
must also be unique ignoring ASCII case.
Each header_env reference must be nonblank, have no surrounding whitespace,
and contain neither = nor NUL. Its environment value must be set, nonblank,
contain no leading or trailing whitespace, be valid Unicode, and be a valid HTTP
header value. Validation errors name the header and environment variable but do
not include the resolved value.
Relay supplies resolved values only as outbound OTLP request headers; it does
not copy them into Event data, OpenTelemetry payloads, resource attributes, or
runtime diagnostics.
The log and metric equivalents are OpenTelemetryLogConfig with
OpenTelemetryLogSubscriber, and OpenTelemetryMetricConfig with
OpenTelemetryMetricSubscriber. Each config takes one required endpoint and
exposes the signal settings documented above. Bare OTLP/HTTP authorities gain
the corresponding standard signal path. Rust, Python, and Node.js expose the
same three independently managed subscriber kinds. The C FFI remains
experimental and source-first.
Direct construction creates one independently managed exporter. Register the
subscriber before instrumented work. During graceful teardown, deregister it,
call the binding’s force-flush method (force_flush() or forceFlush()), and
then call shutdown(). Force flush first crosses Relay’s subscriber barrier
and then flushes the provider. Log shutdown drains the batch queue. Metric
shutdown performs the reader’s final collection; it does not add a second
metric flush. For direct trace and log subscribers, a successful force flush
updates runtime_diagnostics() with any batch queue drops observed so far; the
diagnostic count remains cumulative through later flushes and shutdown.
Log and Metric Subscriber Lifecycle
The following examples create and register direct log and metric subscribers, inspect runtime diagnostics, and perform graceful teardown.
Python
Node.js
Rust
Every direct trace, log, and metric subscriber exposes a bounded runtime
diagnostics snapshot: runtime_diagnostics() in Rust and Python,
runtimeDiagnostics() in Node.js. It reports each runtime condition’s stable
code, occurrence count, and most recent message. C FFI callers use
nemo_relay_otel_subscriber_runtime_diagnostics_json,
nemo_relay_otel_log_subscriber_runtime_diagnostics_json, or
nemo_relay_otel_metric_subscriber_runtime_diagnostics_json. Each writes a
caller-owned bounded JSON array of diagnostic entries that the caller must
release with nemo_relay_string_free. Use diagnostics to monitor rejected
metric marks, capacity limits, and delivery failures without configuring the
observability plugin. The plugin continues to include the same conditions in
its runtime report.
Migrating from Version 3 to Version 4
Version 4 adds the sibling opentelemetry.logs and
opentelemetry.metrics sections. For the complete upgrade path and
programmatic configuration, see
Migrating from Version 3 to Version 4.
Version 2 to Version 3
Version 3 replaces the separate version-2 sections:
- Move the old
opentelemetryfields into one endpoint withtype = "full". - Move the old
openinferencefields into the same section withtype = "openinference". - Use
type = "gen_ai"for standardized GenAI-only output.
Version-2 OTLP section shapes are rejected when version = 3; NeMo Relay does
not silently normalize them. For complete before-and-after configuration and
binding API changes, refer to
Observability Configuration.