OpenTelemetry

View as Markdown

Use the opentelemetry section to configure independent OpenTelemetry Protocol (OTLP) trace, log, and metric pipelines. OpenTelemetry support is always included; no Cargo feature enables or disables it.

Relay classifies each sanitized event before OTLP signal fan-out:

SignalRelay InputOTLP Destination
TracesScope lifecycles and projection-specific non-metric mark handling/v1/traces
LogsNon-metric marks/v1/logs
MetricsMarks with the Relay-owned nemo.relay.metric_measurements schema/v1/metrics

Scope events maintain lineage for log correlation, but Relay does not export them as log records or metric measurements. Metric marks do not fall back to logs or traces when metric export is disabled. A valid mark with the reserved metric schema routes exclusively to metrics. An invalid reserved metric mark reaches no OTLP signal and produces a rate-limited operational diagnostic. Other marks retain the existing projection-specific trace behavior and can be exported as logs.

The metric schema is a NeMo Relay routing contract, not an ATOF-wide metric semantic. ATOF remains at version 0.1, treats data_schema as opaque, and defaults mark data_schema to null. A future specification discussion can evaluate standardizing metric routing, timestamps, aggregation, and exemplars.

Trace Projections

Each endpoint selects one fixed semantic projection:

TypeProjection
fullComplete NeMo Relay projection, including nemo_relay.* attributes and native handling for non-metric marks.
gen_aiOpenTelemetry GenAI semantic conventions only.
openinferenceOpenInference-compatible spans with the existing default handling for non-metric marks.

You can repeat a type or combine types. Each endpoint owns an independent exporter and can use a different endpoint.

The gen_ai projection keeps the OpenTelemetry GenAI semantic-conventions v1.42-era snapshot as its core compatibility baseline. Relay also implements selected newer Development attributes. Refer to GenAI Projection Attribute Support for the exact current inventory and its newer registry baseline.

NeMo Relay uses the currently vendored OpenTelemetry Rust SDK (0.32). It deterministically derives compliant trace and span IDs from Relay lifecycle UUIDs, so endpoints that receive the same event stream use the same identifiers and parentage. Different endpoint types must therefore use independent OTLP destinations; configuring them with the same endpoint and transport is rejected to prevent identifier collisions at the receiver. Duplicate detection compares canonical destinations: HTTP and HTTPS default ports are realized, repeated and trailing path slashes are normalized, and standardized loopback hosts such as localhost, names under .localhost, 127.0.0.0/8, and ::1 are equivalent. Relay does not use DNS resolution for this comparison, and query strings remain significant. Log lineage follows the same root selection as traces. When a local parent was not observed by the exporters, the observed scope starts a trace using its own UUID; a propagation-root UUID alone does not establish an imported parent. Marks under that scope, including late marks within the completed-context retention window, use its trace identity. An explicit imported parent continues the propagated trace. Log-to-span joins require both exporters to observe the corresponding scope lifecycle. Relay propagation continues the Relay-derived trace across the import boundary by default. Use capture_rootless_propagation_context() only when a receiver must start a new OpenTelemetry trace. Carry W3C traceparent and tracestate alongside Relay propagation when an integration also needs to preserve upstream OpenTelemetry sampling or vendor state.

plugins.toml Example

The following version-4 configuration exports a gen_ai trace and derives log and metric destinations from the trace destination.

version = 1
[[components]]
kind = "observability"
enabled = true
[components.config]
version = 4
[components.config.opentelemetry]
enabled = true
[[components.config.opentelemetry.endpoints]]
type = "gen_ai"
endpoint = "http://localhost:4318/v1/traces"
transport = "http_binary"
service_name = "agent-service"
promote_metadata_prefixes = ["app."]
promote_resource_metadata_prefixes = ["deployment."]
max_queue_size = 4096
max_export_batch_size = 512
scheduled_delay_millis = 1000
completed_span_context_ttl_millis = 60000
[components.config.opentelemetry.endpoints.header_env]
authorization = "OTEL_AUTHORIZATION"
[components.config.opentelemetry.endpoints.resource_attributes]
"nv.project" = "observability-dev"
[components.config.opentelemetry.logs]
enabled = true
completed_span_context_ttl_millis = 60000
minimum_severity = "info"
[components.config.opentelemetry.metrics]
enabled = true
temporality = "cumulative"

Relay replaces the terminal /v1/traces with /v1/logs and /v1/metrics. Derived destinations copy the trace endpoint’s transport, authentication, resource attributes, resource-metadata promotion prefixes, service identity, instrumentation scope, and timeout. In this example, all three signals use the same authorization and nv.project routing value.

When enabled = true, configure at least one trace endpoint or an enabled signal with explicit endpoints. NeMo Relay constructs every endpoint before registering subscribers. An invalid endpoint is skipped with an activation warning while valid endpoints register. Activation fails when no trace, log, or metric endpoint can be registered. A delivery failure from one exporter does not stop application work or delivery to the other exporters.

Trace Endpoint Fields

All trace projections (full, gen_ai, and openinference) require HTTPS for remote collectors, with either HTTP or gRPC transport. HTTP is permitted only for localhost and loopback IP addresses such as 127.0.0.1 and [::1]. Redaction does not disable this requirement. Existing remote plaintext trace endpoints must migrate to HTTPS or use a loopback collector.

OTLP/HTTP trace exporters do not follow redirects, even to another HTTPS endpoint. This keeps an export from bypassing transport checks or sending its payload and credentials to another destination. Use the final collector URL. Log and metric HTTP exporters do the same when any header source is configured. gRPC does not use HTTP redirects.

FieldDefaultNotes
typeRequiredfull, gen_ai, or openinference.
endpointRequiredNonblank OTLP endpoint. For OTLP/HTTP, Relay appends /v1/traces when the endpoint contains only a scheme, host, and optional port with no explicit path. Add a trailing / to export traces to the root path. Any other explicit path is preserved. gRPC endpoints are always preserved.
transporthttp_binaryhttp_binary or grpc.
service_nameunknown_serviceservice.name resource attribute.
service_namespaceOmittedOptional service.namespace.
service_versionOmittedOptional service.version.
instrumentation_scopeopentelemetryInstrumentation scope name.
timeout_millis3000OTLP request timeout.
max_queue_sizeEnvironment or 2048Maximum completed spans buffered before this endpoint drops new spans.
max_export_batch_sizeEnvironment or 512Maximum spans exported in one batch; capped at the effective queue size.
scheduled_delay_millisEnvironment or 5000 msMaximum delay before this endpoint exports a non-full batch. A full batch exports sooner.
completed_span_context_ttl_millis60000Positive duration for retaining completed scopes’ trace context for late marks.
headers{}String-to-string exporter headers.
header_env{}Header names mapped to environment variable names containing secret values.
header_file{}Header names mapped to files read immediately before each OTLP HTTP request or gRPC call.
resource_attributes{}String-to-string resource attributes. Relay automatically adds telemetry.sdk.name, telemetry.sdk.language, and telemetry.sdk.version; do not configure those reserved keys.
mark_projectioninheritMark representation for full and openinference: inherit, event, or tool.
mark_exclude_names["llm.chunk"]Mark names excluded from full and openinference projection.
attribute_mappings[]{ key, alias } copies applied by full and openinference projection.
promote_metadata_prefixes[]Literal prefixes that select sanitized Event metadata to copy to top-level span attributes.
promote_resource_metadata_prefixes[]Literal prefixes that select root Scope-start metadata to copy to OTLP resource attributes. Each unique effective resource retains an exporter pipeline for the subscriber lifetime.

Event Metadata Promotion

Wrapped CLI events can include trusted agent_version metadata. To export this generic field as a span attribute, include "agent_version" in promote_metadata_prefixes. Refer to Coding-Agent Identity Metadata for its production and lifecycle semantics.

Set promote_metadata_prefixes on a trace endpoint to copy selected keys from the final sanitized Event metadata into that endpoint’s OpenTelemetry output. The setting defaults to an empty list, so Relay does not promote metadata unless you configure at least one prefix.

BehaviorContract
Promotion prefixesASCII letters, numbers, underscores, and hyphens in nonempty segments separated by single dots, with an optional trailing dot. Matching is literal and case-sensitive.
Supported valuesStrings, booleans, signed 64-bit integers, floating-point numbers, empty arrays, and homogeneous arrays containing one supported primitive type.
Rejected valuesNulls, objects, nested arrays, mixed-type arrays, and integers outside the signed 64-bit range.
Scope lifecycleA Scope-end key is authoritative when Relay constructs the final span. When Scope-end omits the key, the Scope-start value remains.
OpenTelemetry collisionsProjection-owned attributes always win. For full and openinference, configured attribute-mapping aliases also win over promoted metadata.

For example, "app.", "app_", and "app-" are valid literal prefixes. Leading or repeated dots, whitespace, other punctuation, and glob expressions such as "app.*" are rejected.

Matching is case-sensitive and compares the beginning of each key literally. Relay does not infer a dot or metadata-key segment boundary. For example, "app." selects app.name and app.version, but not app_name. The broader "app" prefix selects all three keys. Configure the narrowest prefix that selects the metadata you intend to export.

Scope-start and Scope-end are separate Event records. When Scope-end completes the span, a metadata key present on that Event replaces the corresponding promoted Scope-start value. Mark metadata is promoted to the attributes of the projected span event or tool span. The gen_ai projection continues to omit Marks.

Promotion supports strings, booleans, signed 64-bit integers, floating-point numbers, empty arrays, and homogeneous arrays of those primitive types. Relay omits rejected values and records one bounded runtime diagnostic per rejected key. The diagnostic code is otel.metadata_promotion_value_unsupported.<metadata-key>, its message contains the key and rejection reason, and its count is the number of occurrences for that key. Relay does not record the rejected value or stop trace export. Match the otel.metadata_promotion_value_unsupported. prefix to monitor rejected span metadata keys.

Projection-owned attributes take precedence over promoted metadata with the same key. For full and openinference, configured attribute-mapping aliases also take precedence. Relay also omits selected keys in namespaces owned by Relay or supported semantic projections: nemo_relay., gen_ai., error., exception., input., output., llm., openinference., server., service., session., tool., tool_call., and user.. Relay omits the bare metadata key as well. Rejected keys produce a rate-limited operational diagnostic without dropping the Event or span.

Promotion does not modify the Event or ATOF payload. In OTLP trace output only, Relay removes successfully promoted keys from serialized Relay metadata attributes. This applies to every Scope-start, Scope-end, and Mark event, and also to OpenInference’s metadata JSON attribute. Resource-promotion prefixes participate in that filtering for every event even though only a trace root’s Scope-start metadata can create resource attributes. Keys that cannot be promoted, including values overridden by configured resource attributes, remain in serialized metadata. Use resource_attributes instead for static values that must be attached to every span from an endpoint.

Root Resource Metadata Promotion

Set promote_resource_metadata_prefixes on a trace, log, or metric endpoint to derive resource attributes from the sanitized metadata on a root Scope-start Event. Every span, scoped log, and scoped metric measurement in that scope tree uses the same resource for its configured endpoint; child metadata and later root metadata changes do not modify it. Derived log and metric endpoints inherit the trace endpoint’s prefixes. The prefixes and supported value types match promote_metadata_prefixes.

For explicit log and metric endpoints, the Python plugin helper OpenTelemetrySignalEndpointConfig and the Node.js plugin helper openTelemetrySignalEndpoint accept promote_resource_metadata_prefixes. The Go plugin helper ObservabilityOpenTelemetrySignalEndpointConfig exposes PromoteResourceMetadataPrefixes. The Go trace endpoint helper ObservabilityOpenTelemetryEndpointConfig exposes the same field, which derived log and metric endpoints inherit. Omit the setting to keep promotion disabled. These plugin settings also apply when endpoints are configured through TOML or JSON. The Python, Node.js, and Go direct log and metric subscriber APIs do not expose this option.

Configured service_name, service_namespace, service_version, and explicit resource_attributes take precedence over promoted metadata with the same key. Each distinct effective resource creates a retained OTLP provider pipeline for each enabled signal, so use only controlled, low-cardinality values such as deployment, region, client version, or environment identity. Do not promote request, tenant, or user identifiers.

Each log or metric endpoint retains at most 16 dynamic resource pipelines, plus its configured base provider. Existing resource keys continue to reuse their pipelines at the limit; pipelines are not evicted. Additional resource keys use the base resource and produce otel.resource_metadata_pipeline_limit. This preserves delivery but omits the promoted resource distinction, so metric measurements that differ only by that resource can be aggregated together. The limit is fixed, applies independently to each log and metric endpoint, and does not cap trace pipelines or total process memory.

Each log or metric endpoint tracks at most 4,096 active resource scopes, including nested scopes. Once full, it keeps existing routes and declines new scope starts with otel.resource_metadata_active_scope_limit; marks attached to those untracked scopes use the base resource. Ending a tracked scope frees a slot for a subsequent scope start. Active routes are not evicted by age, so long-running scopes retain their identity. Missing end events can occupy slots until the subscriber is replaced, but cannot grow this active-route map beyond the limit. This limit does not bound completed-route caches, trace correlation state, or total process memory.

Completed resource routes expire before processing an event beyond the TTL; events exactly at the boundary remain linked. Logs use opentelemetry.logs.completed_span_context_ttl_millis (default: 60 seconds). Metrics use a fixed 60-second resource-route TTL. Expiration is driven by event timestamps and runs when another event arrives, rather than on a background timer.

Relay records bounded runtime diagnostics for this setting. A rejected value produces otel.resource_metadata_promotion_value_unsupported.<metadata-key>. A failed resource-pipeline construction produces otel.resource_metadata_pipeline_build_failed; Relay then exports the affected signal through the endpoint’s configured resource instead. Match both the otel.metadata_promotion_value_unsupported. and otel.resource_metadata_promotion_value_unsupported. prefixes to monitor all rejected metadata keys.

Log and Metric Endpoint Resolution

An enabled logs or metrics section can omit endpoints. Relay then derives one signal endpoint from every trace endpoint:

  • A bare HTTP authority, with or without a root trailing /, gains /v1/logs or /v1/metrics.
  • A terminal /v1/traces, including one below a path prefix, is replaced with the signal path. Query parameters are preserved.
  • A gRPC endpoint reuses its authority without path rewriting.
  • A trace endpoint with any other custom path cannot be derived. Configure signal endpoints explicitly in that case.

An explicit nonempty signal endpoint list replaces derivation. Relay preserves an explicit custom signal path exactly, but rejects an obvious standard path for another signal, such as /v1/traces in a log endpoint. An explicit empty list is invalid when the signal is enabled.

The following example sends logs to a custom intake path while metrics continue to derive from the trace endpoint:

[components.config.opentelemetry.logs]
enabled = true
[[components.config.opentelemetry.logs.endpoints]]
endpoint = "https://collector.example/custom/log-intake"
transport = "http_binary"
service_name = "agent-service"
[components.config.opentelemetry.metrics]
enabled = true

The signal endpoint fields are endpoint, transport, headers, header_env, header_file, resource_attributes, promote_resource_metadata_prefixes, service_name, service_namespace, service_version, instrumentation_scope, and timeout_millis. Their defaults match the corresponding trace fields. Each signal rejects duplicate destinations within that signal. Logs, metrics, and traces can share the same authority because OTLP treats them as different signals. All three signals automatically include the reserved telemetry.sdk.name, telemetry.sdk.language, and telemetry.sdk.version resource attributes; configuring those keys in resource_attributes rejects the endpoint.

Trace Batch Processor Configuration

Configure batch processing independently on each endpoint with max_queue_size, max_export_batch_size, and scheduled_delay_millis. When an endpoint omits a field, the corresponding standard OpenTelemetry environment variable applies process-wide. If neither is set, the SDK default applies.

The precedence for each setting is endpoint value, then environment variable, then SDK default. Set environment variables before the plugin activates.

VariableDefaultNotes
OTEL_BSP_MAX_QUEUE_SIZE2048Maximum completed spans buffered per endpoint.
OTEL_BSP_MAX_EXPORT_BATCH_SIZE512Maximum spans exported in one batch; capped at the queue size.
OTEL_BSP_SCHEDULE_DELAY5000 msMaximum delay before exporting a non-full batch.

Endpoint values must be positive integers. If both endpoint size fields are set, max_export_batch_size must not exceed max_queue_size. When one size is inherited, the SDK caps the effective batch size at the effective queue size. The SDK falls back to its default for malformed environment values. Queue and batch sizes count spans, not bytes.

Completed Scope Lineage Retention

Trace endpoints retain a completed scope’s trace and parent span context for completed_span_context_ttl_millis after its scope-end event. A late mark in that window remains attached to the original trace and parent span. When the TTL expires, Relay emits subsequent marks as orphan spans and records the otel.completed_span_context_expired runtime diagnostic when it purges expired contexts.

Keep closed scope handles only for short deferred follow-up work. Prefer emitting an event before the scope closes, or create a new active scope for later work. Increasing the TTL retains more completed contexts: memory grows with the completed-scope rate, TTL, and number of configured trace endpoints.

OTLP logs use the same TTL-based completed-scope lineage behavior. Configure opentelemetry.logs.completed_span_context_ttl_millis independently when log export is enabled. Metrics retain completed resource routes for 60 seconds when resource promotion is enabled, but do not retain trace-correlation context. Do not rely on a closed scope handle for long-running follow-up work; emit the mark before closing the scope or use an active/new scope instead.

If an endpoint’s explicit batch settings are invalid, Relay skips that endpoint and records an observability.invalid_otel_endpoint configuration warning with its opentelemetry.endpoints[N] field. Other valid endpoints continue to activate. Activation still fails when no trace, log, or metric endpoint can be registered.

Relay also skips and logs any trace endpoint that fails during exporter construction. This includes malformed collector destinations that cannot be detected during configuration validation.

Relay’s thread-based batch processor exports serially, so it does not expose the SDK’s concurrent-export setting. It also does not expose a separate batch processor export timeout; use the endpoint’s timeout_millis to bound each OTLP request.

Known limitation: OTLP partial success is not reported. A collector can return a successful OTLP response while rejecting individual spans, log records, or metric data points. With the vendored OpenTelemetry exporter, Relay treats that response as successful: it does not add a runtime diagnostic, and force_flush() and shutdown() can succeed. Monitor collector-side logs and rejection metrics when investigating missing telemetry.

A full queue drops completed spans instead of applying backpressure to application work. Bursts can therefore drop spans that finish late, including an enclosing root span, and leave an incomplete trace in the backend.

The OpenTelemetry SDK logs BatchSpanProcessor.SpanDroppingStarted at warning level when each endpoint first drops a span. It suppresses additional first-drop warnings for that endpoint to avoid a log storm. During graceful shutdown, it logs BatchSpanProcessor.SpansDropped with the endpoint processor’s exact dropped_span_count and max_queue_size.

For plugin-managed exporters, NeMo Relay also records otel.spans_dropped in the active plugin report’s runtime_diagnostics. Its count is the exact number of dropped spans, field identifies the affected opentelemetry.traces[N].endpoint, and message includes the configured endpoint origin (scheme, host, and port), without URL credentials, paths, query parameters, or fragments. If spans were dropped, clearing the plugin returns a delivery failure error and retains the diagnostic for inspection. This error does not disable later plugin configuration.

Increasing an endpoint’s max_queue_size, or the process-wide OTEL_BSP_MAX_QUEUE_SIZE fallback, can reduce the risk for a known burst size, but a finite queue does not guarantee lossless telemetry. Always clear the plugin during graceful shutdown so NeMo Relay can record the final drop count and the SDK can attempt to export queued spans.

Endpoint Capacity and Sizing

NeMo Relay does not impose a maximum number of OpenTelemetry endpoints. Each trace, log, or metric endpoint owns an exporter, signal provider, processor or reader, and exporter runtime resources. Nothing in those export stacks is shared between endpoints. Queue capacity and memory are per endpoint, and total export traffic grows with the endpoint count.

Typical deployments need one to three endpoints. Validate configurations with tens or hundreds of endpoints against the process limits for threads, memory, and network egress before deploying them.

Use header_env for secrets so configuration files contain only environment variable names. Each variable contains the complete header value. NeMo Relay validates variable names without reading their values, then resolves and snapshots the values when the plugin activates. Every referenced variable name must be nonblank and have no surrounding whitespace. Its value must be set and nonblank, with no surrounding whitespace, when the component activates. A header name cannot appear in both headers and header_env, including names that differ only by ASCII case. Reactivate the plugin to pick up a changed environment value.

Use header_file when another process updates a credential in place, such as a projected token file. Each file must be regular and exist when the plugin activates. Relay reads the file only when it exports. Relay removes trailing whitespace. A missing, unreadable, blank, or invalid value fails only that export, and diagnostics do not show the value. A header can use only one of headers, header_env, and header_file. Header names are case-insensitive. This applies to trace, log, and metric endpoints.

Remote trace endpoints must use HTTPS. When you set any header source (headers, header_env, or header_file), remote log and metric endpoints must also use HTTPS. HTTP is allowed only for localhost and loopback IP addresses. OTLP/HTTP exporters do not follow redirects when headers are set. For example, an OIDC token-file writer must write the complete value Bearer eyJ... to the file. Relay does not add the authentication scheme for you.

Process-global OTEL_EXPORTER_OTLP_HEADERS, OTEL_EXPORTER_OTLP_TRACES_HEADERS, OTEL_EXPORTER_OTLP_LOGS_HEADERS, and OTEL_EXPORTER_OTLP_METRICS_HEADERS are rejected because they cannot be isolated between endpoints. Put non-secret values in each endpoint’s headers map and secret variable references in header_env.

full and openinference endpoints retain the legacy mark and attribute-alias controls shown above. gen_ai is standards-only: it ignores those controls and does not emit Relay-private attributes. semantic_selector and capture_content are unsupported.

On a successful tool end span with a present, non-null annotation, the full and openinference projections emit the opaque value as one JSON string attribute named nemo_relay.tool.result.annotation. Relay does not flatten the annotation’s application-defined keys. The gen_ai projection omits this Relay-private attribute.

Emit Log and Metric Marks

Rust, Python, and Node.js generic mark APIs accept optional data_schema and severity values. Prefer the typed metric helper rather than constructing the reserved schema by hand: metric in Rust, Python, and Node.js. The helper validates the complete measurement group before publishing the mark. The exporter validates the sanitized payload again before recording any measurement.

The following examples emit one warning log mark and one metric mark:

from nemo_relay import LogSeverity, MetricKind, MetricMeasurement, MetricValueType
from nemo_relay import scope
scope.event(
"budget-nearly-exhausted",
data={"remaining_tokens": 128},
severity=LogSeverity.Warn,
)
scope.metric(
"tokenomics",
[
MetricMeasurement(
"example.tokens.saved",
MetricKind.Counter,
MetricValueType.U64,
42,
unit="{token}",
description="Tokens avoided",
attributes={"model": "example-model"},
)
],
)

Mark sanitizers run for both calls. Routing uses the immutable data_schema after sanitization, and a metric mark never falls back to the log pipeline.

Log Export

The log pipeline exports one OTLP LogRecord for each sanitized non-metric mark. Marks with data_schema = null and marks with an application-defined schema are logs. Any mark that uses the reserved nemo.relay.metric_measurements schema name is routed away from logs, including unsupported schema versions and invalid metric payloads. Scope start and end events update the lineage used for correlation but do not become log records.

Use the typed severity argument on the generic mark API. Relay stores it in sanitizer-visible metadata as nemo_relay.log.severity. The typed argument overrides that metadata key and requires metadata to be an object. After mark sanitizers run, Relay parses the remaining key, defaults an absent key to info, and drops a log record with an invalid value. Supported values are trace, debug, info, warn, and error; warning is accepted as an alias for warn.

The logs section applies these processing settings to every log endpoint:

FieldDefaultNotes
minimum_severityinfoIndependent telemetry-log threshold. It does not inherit process logging settings.
max_queue_size2048Maximum queued log records.
max_export_batch_size512Maximum records in one batch; must not exceed the queue size.
scheduled_delay_millis1000Maximum delay before exporting a partial batch.
completed_span_context_ttl_millis60000Positive duration for retaining completed scope parent context for late logs. Contexts exactly at the TTL boundary remain linked.

Relay maps the event timestamp to the log timestamp and post-sanitization processing time to the observed timestamp. Sanitized data becomes the structured body. An absent or top-level JSON null payload has no body; a nested JSON null becomes the string "null" because OTLP AnyValue has no null variant.

The log attributes preserve the mark name, UUID, optional parent UUID, category and category profile, schema, sanitized metadata, and ATOF version under nemo_relay.* keys. A mark in a resolvable active or completed scope receives trace and span context. An orphan mark receives no invented trace context. Relay leaves OTLP event_name unset with the currently vendored OpenTelemetry SDK (0.32) and retains the dynamic name in nemo_relay.mark.name.

Telemetry logs are separate from NeMo Relay’s operational stderr and file logging. minimum_severity does not inherit NEMO_RELAY_LOG, and operational diagnostics are not fed back into ATOF or OTLP.

Metric Export

The metric pipeline consumes only sanitized marks with this exact schema:

{
"data_schema": {
"name": "nemo.relay.metric_measurements",
"version": "1"
},
"data": {
"measurements": [
{
"name": "example.tokens.saved",
"kind": "counter",
"value_type": "u64",
"value": 42,
"unit": "{token}",
"description": "Tokens avoided",
"attributes": {"model": "example-model"},
"boundaries": null
}
]
}
}

A mark with the reserved schema name and an unsupported version or invalid payload is dropped from both logs and metrics. Relay emits a rate-limited operational diagnostic without creating another ATOF or OTLP event.

The measurements array is required and nonempty, and unknown fields are rejected. Each measurement is an SDK recording operation, not a pre-aggregated OTLP point:

KindAllowed value_typeRecording Semantics
counteru64, or nonnegative finite f64Addition
up_down_counteri64, or finite f64Signed addition
gaugeu64, i64, or finite f64Current value
histogramu64, or finite f64Distribution sample

Unsigned values must not exceed i64::MAX, which prevents loss in the pinned OTLP conversion. Metric names must be 1 to 255 ASCII bytes, start with a letter, and contain only letters, digits, _, ., -, or /. Units must be ASCII and at most 63 bytes. Optional histogram boundaries can include negative values, but every boundary must be finite, strictly increasing, unique, and the list can contain at most 64 entries.

Attributes can contain strings, Booleans, signed integers, finite doubles, and homogeneous nonempty arrays of those primitive types. Blank keys, nulls, nested objects, mixed arrays, and unsigned integers above i64::MAX are invalid. Do not use event UUIDs, timestamps, metadata, trace IDs, or other high-cardinality values as metric attributes.

Relay treats the complete mark atomically. It records no measurements from the mark when a measurement is invalid, an instrument descriptor conflicts, or an instrument limit would be exceeded. Instrument names compare case-insensitively, and a name must retain its kind, numeric type, unit, description, and histogram boundaries for the lifetime of that destination.

The metrics section applies these settings to every metric endpoint:

FieldDefaultNotes
export_interval_millis60000Periodic collection interval.
temporalitycumulativecumulative, delta, or low_memory.
max_instruments256Retained instrument descriptors per destination.
cardinality_limit2000SDK series limit per instrument.

Metric points use SDK collection timestamps. The currently vendored OpenTelemetry SDK (0.32) cannot preserve the source mark timestamp or attach a trace-linked exemplar through this path. Relay does not emulate correlation with high-cardinality attributes.

GenAI Projection

Set an OpenTelemetry endpoint’s type to gen_ai to select this projection:

[[components.config.opentelemetry.endpoints]]
type = "gen_ai"
endpoint = "http://localhost:4318/v1/traces"

The gen_ai endpoint uses these operation names:

Relay scopeOpenTelemetry operation
Agentinvoke_agent
LLMchat, generate_content, or text_completion
Toolexecute_tool
Embedderembeddings
Retrieverretrieval

Marks are omitted. Relay scope types without GenAI semantics are emitted as minimal internal spans so that the original span parentage is preserved. This projection never emits nemo_relay.* fields. LLM spans include the gen_ai.system_instructions, gen_ai.input.messages, and gen_ai.output.messages attributes as JSON strings that follow the OpenTelemetry GenAI schemas. Each attribute is emitted only when its normalized instructions, messages, or response content is present. Redact sensitive content with an LLM or event sanitizer. Tool content is exported by default; retrieval payloads are not exported. Set top-level enable_full_payloads = true to retain complete sanitized LLM request history on every start span.

GenAI Projection Attribute Support

This inventory follows the 72-attribute OpenTelemetry GenAI registry snapshot used for the current Relay audit. Every listed upstream attribute is marked Development. This newer inventory includes attributes added or renamed after Relay’s v1.42-era compatibility baseline.

Supported means Relay’s native GenAI projection emits the exact attribute on an applicable signal when an authoritative normalized or instrumentation source exists. It does not mean every integration supplies that optional source. Partial means Relay emits the attribute but cannot represent the complete upstream cardinality or value-shape contract. Not supported means the native projection does not emit it. The gen_ai.* namespace is reserved, so generic metadata promotion cannot be used to add an unsupported GenAI attribute.

Projection AttributeSupportDescription and Notes
gen_ai.agent.descriptionSupportedApplication-provided free-form agent description. Relay emits it on Agent scopes and marked CLI turns when canonical or recognized alias metadata supplies it.
gen_ai.agent.idNot supportedStable identifier of a hosted GenAI agent resource. Relay does not yet distinguish hosted-agent create or client operations, and it does not substitute transient scope, session, subagent, or harness identifiers.
gen_ai.agent.nameSupportedApplication-provided human-readable agent name. Relay uses Agent scope identity and explicit applicable turn or tool metadata; marked turns omit it when no authoritative name exists.
gen_ai.agent.versionNot supportedVersion of a hosted GenAI agent. Relay does not yet classify hosted-agent operations. The generic agent_version captured by wrapped CLI launches identifies the harness executable and is intentionally not projected as this attribute.
gen_ai.conversation.compactedNot supportedIndicates that the effective context is a compacted view of an earlier conversation. Relay does not yet propagate positive compaction state into the later LLM operation; false must not be emitted.
gen_ai.conversation.idSupportedStable conversation, session, or thread identifier. Relay projects the canonical key or recognized conversation, session, and thread aliases on applicable Agent, turn, LLM, and tool spans.
gen_ai.data_source.idSupportedIdentifier of the GenAI data source. Relay projects it on retriever spans from the canonical key or recognized data-source aliases; prefer the GenAI system identifier over an external storage name.
gen_ai.embeddings.dimension.countSupportedRequested output embedding dimension count. Relay emits a positive integer on embedder spans from the canonical key or dimensions.
gen_ai.evaluation.explanationNot supportedEvaluator-provided explanation for an assigned score. The convention places it on a gen_ai.evaluation.result event; Relay has no corresponding standard event projection.
gen_ai.evaluation.nameNot supportedName of the evaluation metric. Relay has no gen_ai.evaluation.result event projection.
gen_ai.evaluation.score.labelNot supportedHuman-readable interpretation of an evaluation score. Relay has no gen_ai.evaluation.result event projection.
gen_ai.evaluation.score.valueNot supportedNumeric evaluation score. Relay has no gen_ai.evaluation.result event projection.
gen_ai.input.messagesSupportedChat history supplied to the model. Relay serializes retained normalized and sanitized messages as schema-shaped JSON text. enable_full_payloads controls complete request-history retention, not whether present content can be projected.
gen_ai.memory.query.textNot supportedSearch query used to retrieve memories. Relay has no standard memory-operation lifecycle or sensitive-content opt-in contract for this value.
gen_ai.memory.record.countNot supportedNumber of memory records relevant to an operation. Relay has no normalized memory-result lifecycle.
gen_ai.memory.record.idNot supportedUnique memory-record identifier. Relay has no normalized memory-record identity contract.
gen_ai.memory.recordsNot supportedMemory records stored or retrieved by a memory operation. Relay has no memory lifecycle, normalized record shape, or sensitive-content opt-in contract.
gen_ai.memory.store.idNot supportedUnique identifier of a memory store. Relay has no standard memory scope or operation contract.
gen_ai.operation.nameSupportedName of the GenAI operation. Relay projects invoke_agent, chat, generate_content, text_completion, execute_tool, embeddings, or retrieval on the corresponding scopes and marked CLI turn roots.
gen_ai.output.messagesPartialModel output messages, with one message per returned choice or candidate. Relay currently normalizes and projects only one assistant candidate, so additional parallel generations are not retained.
gen_ai.output.typeSupportedOutput modality requested by the client. Exact canonical metadata wins; Relay otherwise derives json, text, or speech from supported OpenAI request shapes and omits ambiguous, multiple, or unsupported modalities.
gen_ai.prompt.nameNot supportedName that uniquely identifies a prompt template. Relay has no authoritative cross-provider prompt identity source.
gen_ai.prompt.variable.<name>Not supportedRuntime value supplied for a named prompt-template variable. Relay has no normalized variable map or explicit sensitive-content policy for these dynamic attributes.
gen_ai.prompt.versionNot supportedVersion of the prompt template. Relay has no authoritative cross-provider prompt-version source.
gen_ai.provider.nameSupportedGenAI provider identified by the instrumentation. Relay uses explicit canonical or alias metadata, recognized routes and event names, or the normalized provider API; it can emit custom values such as oci.genai.
gen_ai.request.choice.countSupportedRequested number of candidate completions. Relay projects non-default OpenAI Chat n values; the default value of one is omitted.
gen_ai.request.encoding_formatsSupportedRequested embedding encoding formats. Relay projects the canonical key or recognized singular and plural aliases on embedder spans.
gen_ai.request.frequency_penaltySupportedRequest frequency-penalty setting. Relay projects it from normalized OpenAI Chat requests.
gen_ai.request.max_tokensSupportedMaximum number of tokens requested for generation. Relay projects the normalized request value.
gen_ai.request.modelSupportedName of the model targeted by the request. Relay projects normalized request or authoritative model metadata on applicable spans.
gen_ai.request.presence_penaltySupportedRequest presence-penalty setting. Relay projects it from normalized OpenAI Chat requests.
gen_ai.request.previous_response.idSupportedIdentifier of a prior response used as context for the current operation. Relay projects the normalized previous-response ID. This attribute is newer than the v1.42-era baseline.
gen_ai.request.reasoning.levelSupportedRequested reasoning or thinking effort. Relay uses exact canonical metadata, OpenAI Chat reasoning_effort, or normalized reasoning effort.
gen_ai.request.seedSupportedSeed intended to make repeated requests more deterministic. Relay projects it from normalized OpenAI Chat requests.
gen_ai.request.stop_sequencesSupportedSequences that stop further token generation. Relay projects the normalized request list.
gen_ai.request.streamSupportedIndicates a streaming request. Relay emits only true; absence represents a non-streaming request as required by the convention.
gen_ai.request.stream_cursorNot supportedCursor used to resume a streamed response after the last received event. Relay has no fetch or resume operation classification or normalized cursor source.
gen_ai.request.temperatureSupportedRequest temperature setting. Relay projects the normalized request value.
gen_ai.request.top_kSupportedTop-K sampling limit used during generation. Relay projects normalized Anthropic top_k; it does not misclassify OpenAI top_logprobs as this attribute.
gen_ai.request.top_pSupportedRequest nucleus-sampling setting. Relay projects the normalized request value.
gen_ai.response.finish_reasonsPartialOrdered reasons that each returned generation stopped. Relay retains and emits one normalized finish reason, so it cannot represent multiple candidates or an expected generation that ended before producing a normal reason.
gen_ai.response.idSupportedUnique completion or response identifier. Relay projects the normalized response ID.
gen_ai.response.modelSupportedName of the model that generated the response. Relay projects normalized LLM response data or an authoritative embedder response source.
gen_ai.response.statusNot supportedProvider-reported lifecycle status when a response is fetched or polled. Relay does not classify fetch or poll operations and intentionally does not copy ordinary inference status into this attribute.
gen_ai.response.time_to_first_chunkSupportedSeconds from managed request execution to the first received provider protocol chunk. This is not time to first text token. Relay omits it for non-streaming calls and streams with no chunk, and emits the companion standard histogram only when provider identity is known.
gen_ai.retrieval.documentsNot supportedDocuments returned by retrieval. Relay has no normalized document-result source or content policy; raw retrieval payloads are not exported by the GenAI projection.
gen_ai.retrieval.query.textNot supportedQuery text used for retrieval. Relay has no explicit normalized query source or sensitive-content opt-in contract.
gen_ai.retrieval.top_kSupportedMaximum number of requested retrieval documents. Relay projects the canonical key or top_k on retriever spans.
gen_ai.system_instructionsSupportedSystem instructions supplied separately from chat history. Relay serializes retained normalized and sanitized instructions as schema-shaped JSON text when present.
gen_ai.token.typeNot supportedToken category used by the standard gen_ai.client.token.usage metric. It is not a span attribute, and Relay does not automatically emit that metric; custom typed metric marks can carry it independently.
gen_ai.tool.call.argumentsPartialParameters passed to a tool call. Relay exports every non-null sanitized value as canonical JSON text. OpenTelemetry expects object-shaped data, so arrays and scalar values are preserved but do not satisfy the upstream schema.
gen_ai.tool.call.idSupportedTool-call identifier. Relay prefers the typed call ID, then explicit instrumentation metadata, and never derives identity from argument or result payloads.
gen_ai.tool.call.resultPartialSuccessful result returned by a tool call. Relay omits failed results and exports every non-null sanitized value as canonical JSON text. OpenTelemetry expects object-shaped data, so arrays and scalar values do not satisfy the upstream schema.
gen_ai.tool.definitionsSupportedTool definitions available to the agent or model. Relay projects minimal schema-shaped definitions containing required type and name identities while omitting optional descriptions and parameter schemas.
gen_ai.tool.descriptionSupportedTool description. Relay emits nonblank explicit instrumentation metadata and never infers it from invocation arguments.
gen_ai.tool.nameSupportedName of the tool used by the agent. Relay uses explicit tool identity metadata or the tool scope name.
gen_ai.tool.typeSupportedType of tool used by the agent, such as function, extension, or datastore. Relay uses explicit metadata or source-backed harness defaults and preserves explicit refinements.
gen_ai.usage.audio.cache_read.input_tokensNot supportedAudio input tokens served from a provider cache. Relay’s normalized usage contract has no audio and cache modality breakdown.
gen_ai.usage.audio.input_tokensNot supportedAudio input-token count. Relay’s normalized usage contract has no audio modality breakdown.
gen_ai.usage.audio.output_tokensNot supportedAudio output-token count. Relay’s normalized usage contract has no audio modality breakdown.
gen_ai.usage.cache_read.input_tokensSupportedInput tokens served from a provider-managed cache. Relay projects normalized cache-read usage and includes known provider cache accounting in the aggregate input count.
gen_ai.usage.cache_write.input_tokensSupportedInput tokens written to a provider-managed cache. Relay projects normalized cache-write usage. This newer spelling replaces gen_ai.usage.cache_creation.input_tokens from the v1.42-era registry snapshot.
gen_ai.usage.image.cache_read.input_tokensNot supportedImage input tokens served from a provider cache. Relay’s normalized usage contract has no image and cache modality breakdown.
gen_ai.usage.image.input_tokensNot supportedImage input-token count. Relay’s normalized usage contract has no image modality breakdown.
gen_ai.usage.image.output_tokensNot supportedImage output-token count. Relay’s normalized usage contract has no image modality breakdown.
gen_ai.usage.input_tokensSupportedTotal GenAI input-token count. Relay uses normalized usage and includes separately reported provider cache reads and writes when required to form the total.
gen_ai.usage.output_tokensSupportedTotal GenAI output-token count. Relay projects normalized completion usage.
gen_ai.usage.reasoning.output_tokensSupportedOutput tokens used for reasoning or extended thinking. Relay uses exact canonical metadata, OpenAI Responses reasoning details, or Gemini thought-token usage.
gen_ai.usage.text.cache_read.input_tokensNot supportedText input tokens served from a provider cache. Relay’s normalized usage contract has no text and cache modality breakdown.
gen_ai.usage.text.input_tokensNot supportedText input-token count. Relay’s normalized usage contract has no text modality breakdown.
gen_ai.usage.text.output_tokensNot supportedText output-token count. Relay’s normalized usage contract has no text modality breakdown.
gen_ai.workflow.nameNot supportedApplication-provided low-cardinality workflow name. Relay has no Workflow scope or invoke_workflow operation classification.

Usage is recorded per model request. Relay’s normalized response usage retains an optional uncached_input_tokens count only when the provider’s response contract proves it. The standard GenAI export deliberately does not add a nonstandard uncached field: consumers can use the total gen_ai.usage.input_tokens with the separately reported cache-read and cache-write counts. An omitted cache or uncached count means unavailable, not zero.

For managed streaming LLM calls, the end span also includes gen_ai.response.time_to_first_chunk in seconds. Relay measures from the start of managed stream execution until the first provider protocol chunk is received; this is not a first-text-token measurement. The attribute is omitted for non-streaming calls and streams that never yield a chunk.

The metric endpoint records the same sample as the standard gen_ai.client.operation.time_to_first_chunk f64 histogram with unit s when Relay can determine the provider dimension.

The gen_ai projection includes sanitized gen_ai.tool.call.arguments at tool start, gen_ai.tool.call.result on successful completion, and gen_ai.tool.definitions on inference spans by default. These attributes can contain sensitive information. Use the PII redaction trajectory_context preset to remove opaque payloads while retaining analytical structure and trace parentage. Credential removal and event sanitizers run before projection. enable_full_payloads controls LLM request-history retention independently.

Arguments and results are emitted as canonical JSON strings. Relay parses serialized JSON before projection, then preserves every non-null sanitized JSON value, including objects, arrays, strings, booleans, and numbers. Failed tool calls do not emit a result attribute. Result annotations remain separate and are never included in this projection. Tool definitions include only required type and name properties. Known OpenAI Responses built-ins use their native type as identity; recognized Gemini native tool groups emit an identity for each tool in the group. Explicit native names are preserved. Unknown unnamed native definitions are omitted rather than assigned invented identities.

Tool Identity

Tool identity comes from instrumentation metadata and the typed tool-call ID, never from similarly named fields in tool arguments or results. Nonblank string metadata supports gen_ai.tool.name, gen_ai.tool.type / tool_type, gen_ai.tool.call.id / tool_call_id, gen_ai.tool.description / tool_description / description, and gen_ai.agent.name / agent_name. The typed call ID takes precedence. Absent tool names use the scope name; unknown descriptions, types, and agent names are omitted.

Claude Code, Codex, and Pi tool hooks default gen_ai.tool.type to function because these harnesses execute tools locally. They default gen_ai.agent.name to claude-code, codex, or pi, or subagent:<id> when ownership is known. These defaults apply to paired hooks and post-only hooks that synthesize a start, without replacing explicitly supplied nonblank string metadata, including the tool_type and agent_name aliases. Generic gateway sessions do not infer a tool type or root executing-agent name. Invocation arguments are never used to infer tool descriptions.

Harness defaults carry internal provenance. Explicit completion metadata may refine inferred tool type and agent name, including through their aliases. Explicit start metadata and typed tool-call IDs remain authoritative; later inferred values cannot replace them. This provenance is not exported as a GenAI attribute.

Claude Code and Codex MCP tool hooks with a qualified mcp__<server>__<tool> name receive mcp.method.name = "tools/call" in event metadata, including post-only hooks. Relay preserves existing method metadata and does not infer MCP identity for ordinary tools, Pi, or generic gateway events. The original qualified tool name remains unchanged. To include the method in OTLP span attributes, configure promote_metadata_prefixes to include "mcp.method.name".

This identifies a requested MCP tool operation, not proof that a request reached the server or succeeded. Permission-denied calls retain their denial and error metadata. The server segment is a configured alias; Relay does not derive server.address, mcp.session.id, or connection status from it. Hook duration measures the harness tool scope, not necessarily MCP transport latency.

Error Type Mapping

For managed LLM, tool, and stream failures, NeMo Relay maps structured FlowError values to the OpenTelemetry error.type attribute:

Relay errorerror.type
AlreadyExistsalready_exists
NotFoundnot_found
InvalidArgumentinvalid_argument
ScopeStackEmptyscope_stack_empty
GuardrailRejectedguardrail_rejected
Upstream connection failureconnection_error
Upstream timeouttimeout
Upstream retryable statusretryable_status
Upstream context-window failurecontext_window
Upstream model unavailablemodel_unavailable
Upstream authentication failureauthentication
Upstream invalid requestinvalid_request
Other upstream failureupstream_error
Internalinternal_error
Binding callback exceptioninternal_error

External application and callback exceptions that do not have a more specific FlowError classification emit internal_error. Python and JavaScript callback boundaries also preserve the exception class separately, and both the full and gen_ai projections emit an exception span event with exception.type. NeMo Relay does not inspect error messages to recover exception class names. When an errored parent span has no useful classification of its own, it inherits the failed descendant’s error.type and exception type. When no structured FlowError is available, such as a cancellation or dropped execution, the projection emits _OTHER. Caller-provided error.type and exception.type metadata take precedence over values derived from FlowError. An authenticated Claude Code or Codex permission rejection closes the matched active tool span with error.type = "guardrail_rejected". The corresponding permission guardrail events carry the canonical gen_ai.tool.call.id in event metadata so subscribers can correlate the decision with that tool call.

FlowError is an exhaustive Rust enum. Rust callers upgrading to this release must handle the new CallbackException variant in exhaustive matches. It maps to the same internal status as Internal, while retaining exception_type for observability projection.

Direct Subscribers

from nemo_relay import OpenTelemetryConfig, OpenTelemetrySubscriber
config = OpenTelemetryConfig(
"gen_ai",
"http://localhost:4318/v1/traces",
)
config.service_name = "agent-service"
config.header_env = {"authorization": "OTEL_AUTHORIZATION"}
subscriber = OpenTelemetrySubscriber(config)

Set each referenced environment variable before constructing the subscriber. Direct trace, log, and metric configs resolve header_env when the subscriber is constructed and retain that value for the subscriber’s activation. Changing the process environment affects only a subsequently constructed subscriber. Static headers remain unchanged. A header name cannot appear in both maps, including names that differ only by ASCII case, and names within header_env must also be unique ignoring ASCII case.

Each header_env reference must be nonblank, have no surrounding whitespace, and contain neither = nor NUL. Its environment value must be set, nonblank, contain no leading or trailing whitespace, be valid Unicode, and be a valid HTTP header value. Validation errors name the header and environment variable but do not include the resolved value. Relay supplies resolved values only as outbound OTLP request headers; it does not copy them into Event data, OpenTelemetry payloads, resource attributes, or runtime diagnostics.

The log and metric equivalents are OpenTelemetryLogConfig with OpenTelemetryLogSubscriber, and OpenTelemetryMetricConfig with OpenTelemetryMetricSubscriber. Each config takes one required endpoint and exposes the signal settings documented above. Bare OTLP/HTTP authorities gain the corresponding standard signal path. Rust, Python, and Node.js expose the same three independently managed subscriber kinds. The C FFI remains experimental and source-first.

Direct construction creates one independently managed exporter. Register the subscriber before instrumented work. During graceful teardown, deregister it, call the binding’s force-flush method (force_flush() or forceFlush()), and then call shutdown(). Force flush first crosses Relay’s subscriber barrier and then flushes the provider. Log shutdown drains the batch queue. Metric shutdown performs the reader’s final collection; it does not add a second metric flush. For direct trace and log subscribers, a successful force flush updates runtime_diagnostics() with any batch queue drops observed so far; the diagnostic count remains cumulative through later flushes and shutdown.

Log and Metric Subscriber Lifecycle

The following examples create and register direct log and metric subscribers, inspect runtime diagnostics, and perform graceful teardown.

from nemo_relay import (
OpenTelemetryLogConfig,
OpenTelemetryLogSubscriber,
OpenTelemetryMetricConfig,
OpenTelemetryMetricSubscriber,
)
logs = OpenTelemetryLogSubscriber(
OpenTelemetryLogConfig("http://localhost:4318/v1/logs")
)
metrics = OpenTelemetryMetricSubscriber(
OpenTelemetryMetricConfig("http://localhost:4318/v1/metrics")
)
logs.register("otlp-logs")
metrics.register("otlp-metrics")
try:
# Run instrumented work here.
for diagnostic in logs.runtime_diagnostics().entries:
print(diagnostic.code, diagnostic.message)
finally:
logs.deregister("otlp-logs")
logs.force_flush()
logs.shutdown()
metrics.deregister("otlp-metrics")
metrics.force_flush()
metrics.shutdown()

Every direct trace, log, and metric subscriber exposes a bounded runtime diagnostics snapshot: runtime_diagnostics() in Rust and Python, runtimeDiagnostics() in Node.js. It reports each runtime condition’s stable code, occurrence count, and most recent message. C FFI callers use nemo_relay_otel_subscriber_runtime_diagnostics_json, nemo_relay_otel_log_subscriber_runtime_diagnostics_json, or nemo_relay_otel_metric_subscriber_runtime_diagnostics_json. Each writes a caller-owned bounded JSON array of diagnostic entries that the caller must release with nemo_relay_string_free. Use diagnostics to monitor rejected metric marks, capacity limits, and delivery failures without configuring the observability plugin. The plugin continues to include the same conditions in its runtime report.

Migrating from Version 3 to Version 4

Version 4 adds the sibling opentelemetry.logs and opentelemetry.metrics sections. For the complete upgrade path and programmatic configuration, see Migrating from Version 3 to Version 4.

Version 2 to Version 3

Version 3 replaces the separate version-2 sections:

  • Move the old opentelemetry fields into one endpoint with type = "full".
  • Move the old openinference fields into the same section with type = "openinference".
  • Use type = "gen_ai" for standardized GenAI-only output.

Version-2 OTLP section shapes are rejected when version = 3; NeMo Relay does not silently normalize them. For complete before-and-after configuration and binding API changes, refer to Observability Configuration.