Release Notes for NVIDIA NeMo Relay
This page contains the release notes for NVIDIA NeMo Relay.
Release 0.8
NeMo Relay 0.8 strengthens plugin extensibility, managed tool execution, and observability while introducing source-breaking migration steps for native Rust plugins, Node.js streaming intercepts, and tool-result callbacks.
Highlights
The following updates expand native Rust plugin capabilities:
- Typed native Rust middleware can await Tokio timers, I/O, codecs, unary
continuations, and downstream LLM streams on an SDK-owned executor. Native
ABI v4 adds completion-scoped codec operations, pull-based LLM stream
continuations, and
emit_mark_v2support for typed data schemas and log severity. Existing raw ABI v2 and v3 binaries remain loadable without a rebuild for host-table compatibility; rebuild only when the plugin also needs the changed native API 1 tool-result contract. - Tool callbacks and execution-intercept continuations now share the canonical
ToolExecutionResultcontract across Rust, Python, Node.js, Go, the public C surface, native plugins, andgrpc-v1workers. The application-ownedresultcan travel with an optional opaqueannotationfor observability. - ATOF, ATIF, and the
fullandopeninferenceOpenTelemetry projections preserve sanitized tool result annotations without flattening their application-defined schemas. - Managed tool execution accepts an optional provider- or harness-supplied tool-call ID and preserves it on the matching start and end events across success, error, and cancellation paths.
Runtime Controls and Caching
- Conditional middleware guardrails let applications and plugins temporarily exclude matching global registrations, such as an exporter during an outage or an expensive intercept while a dependency is unhealthy. Gates are fail-open, so a failed control callback does not silently remove runtime behavior. See Conditional Middleware Guardrails.
- Event metadata injectors can enrich events from application code and plugins across Rust, Python, Node.js, Go, and the C FFI. Selected metadata keys can also be promoted into OpenTelemetry attributes, making useful application context available to telemetry backends without changing event schemas.
- The Adaptive response cache can now cache explicitly classified, read-only tool results. Tool caching is disabled by default because a cache hit skips the live tool call; configure a stable TTL and tool policy only when that is safe for the tool’s behavior. See Response Cache.
Provider and Framework Support
- Built-in codecs now support OCI Generative AI chat payloads and Gemini
generateContentpayloads, including streaming and normalized lifecycle data. Use these codecs directly when a framework sends either provider’s request shape. See Provider Codecs. - Managed LLM requests now carry a runtime-owned W3C
traceparentheader, allowing downstream provider work to remain connected to the Relay trace. - The Deep Agents integration now models the orchestrator and in-process subagents as nested semantic Agent scopes. This makes parent-child activity easier to understand without turning internal LangGraph nodes into extra agent spans.
OpenTelemetry Logs and Metrics
- Independent OTLP log and metric pipelines are available alongside trace export. Observability configuration version 4 derives log and metric destinations from a trace endpoint when their endpoint lists are omitted.
- Rust, Python, Node.js, and the experimental C FFI expose typed mark schema and severity options plus metric helpers. Direct subscribers and native and worker plugin SDKs expose bounded runtime diagnostics snapshots.
- OpenTelemetry keeps valid endpoints active when a peer endpoint is invalid, retains completed trace context for late events for a configurable period, and reports signal-specific delivery failures without exposing sensitive endpoint details.
Support Matrix and Compatibility Updates
The Support Matrix is the canonical reference for supported platforms and architectures, worker runtimes, coding agents, and integrations. It also records current limitations, including platform-specific worker requirements.
Migration guidance for upgrading from 0.7 to 0.8 is available in the Migration Guides.
Breaking Changes
- Typed native Rust middleware callbacks now return futures and receive owned
arguments. Native subscribers and raw synchronous ABI entry points are
unchanged. Rebuild typed native Rust middleware plugins with
nemo-relay-plugin0.8.0 and declarecompat.relay = ">=0.8.0,<1.0"; see the native Rust migration guide. - Node.js LLM streaming execution intercepts now receive and return lazy async
iterables. Existing intercepts that collect
next(request)into an array or return a scalar or array must use an async generator instead. See the Node.js stream-intercept migration guide. - Relay-created plugin-component effective names now use a versioned format
that includes a one-based component ordinal, including ordinal
1for a singleton component. Effective names are an internal implementation detail: discover them from the active runtime instead of constructing or persisting them. - Managed tool callbacks, execution-intercept continuations, execute-helper
returns, and manual tool end APIs now use
ToolExecutionResultinstead of a raw JSON result. Applications that need the original business payload must read itsresultfield. Refer to the tool result migration guide. - Relay 0.8 resets native API 1 and the
grpc-v1tool-result contract to the canonicalToolExecutionResultshape. Rebuild native API 1 and worker plugins that use tool callbacks, execution intercepts, or manual tool-end APIs, and declare acompat.relayrange that excludes Relay versions before 0.8. Frozen raw ABI v2 and v3 binaries that do not use this changed contract remain loadable without a rebuild. The native ABI v4 table remains unchanged. The worker package and RPC method names remainnemo.relay.worker.v1, but theToolNextresponse and tool-execution outcome field now use structured protobuf messages. Regenerate custom worker bindings before rebuilding. - Repository-local
.nemo-relay/config.toml,plugins.toml, and.dynamic-plugins.jsonfiles are no longer discovered or activated. Runtime resolution is explicit-or-user followed by higher-precedence system policy. - Project setup scopes and the
--projectandconfig --reset --scopeinterfaces have been removed. Setup and reset now target XDG user configuration;--user,--global,--config, and--plugin-config-pathremain supported. nemo-relay doctorreports ignored ancestor project configuration as warning-only migration diagnostics. Relay does not automatically move, rewrite, or delete those legacy files.- Removed Hermes-specific support from the NeMo Relay CLI, including the
hermesshortcut,run --agent hermes, and Hermes-specific install, uninstall, doctor, configuration, MCP selection, and/hooks/hermespaths. NeMo Relay is built into Hermes Agent, and Hermes Agent understands NeMo Relay plugin configurations. No separate observability plugin or Relay CLI setup is required. - Removed the experimental, service-backed Switchyard integration, including
the
nemo-relay-switchyardcrate, CLIswitchyardfeature, built-inswitchyardcomponent, and service examples. Switchyard 0.3.0 will distribute and document the Switchyard-owned dynamic plugin. Refer to the Switchyard migration guide.
Refer to Migration Guides for destination paths and explicit-file alternatives.
Fixed Known Issues in 0.8
- Codex image-generation requests now pass through the local Relay gateway to
the configured OpenAI upstream. Persistent Codex installations must refresh
the
nemo-relay-openaiprovider configuration; see Migration Guides. - Restored a trailing
/on an OTLP/HTTP trace endpoint as an explicit root-path destination. A bare HTTP authority still defaults to/v1/traces; when version 4 implicitly derives log or metric destinations from either form, it uses/v1/logsor/v1/metricsrespectively. Explicit trace, log, and metric endpoint paths remain unchanged. - OpenTelemetry activation now isolates invalid endpoints so they do not disable valid trace, log, or metric destinations. Repeated shutdown is also safe, and blocked trace flushes no longer delay event delivery to other subscribers.
- OTLP diagnostics now retain distinct causes, identify the affected signal, report direct log queue drops and trace or metric export failures, and redact sensitive endpoint details from trace-export failures.
- Successful force flushes for direct OTLP trace and log subscribers now expose all queue drops observed so far in runtime diagnostics. The reported count remains cumulative through later flushes and shutdown.
- Deferred marks can retain their completed trace lineage for a configurable time instead of becoming orphaned after a fixed number of completed scopes.
- ATIF exports now correlate OpenAI Responses function-call results with the
semantic invocation
call_id, so the exported tool observation remains connected to the matching tool call instead of being orphaned by the response item’s separate identifier.
Other Improvements
- Manually observed LLM responses can now receive configured model-pricing estimates when they include sufficient normalized model and usage data; provider-reported costs remain authoritative.
nemo-relay run --dry-runnow warns when the forwarded command repeats the explicitly selected coding-agent executable, helping catch a common launch mistake without changing live-launch behavior.- Set
NEMO_RELAY_PLUGIN_SNAPSHOT_DIRto choose the parent directory for plugin activation snapshots when system temporary paths are unsuitable or too long.
Known Issues in 0.8
- OTLP collectors can return a successful response while rejecting individual spans, log records, or metric data points. Relay 0.8 does not report these partial successes in runtime diagnostics, and flush or shutdown can still succeed. Monitor collector-side rejection metrics and logs; see OpenTelemetry.
- Go and the raw C FFI remain experimental and source-first. Generated API pages focus on Rust, Python, and Node.js.
- Local coding-agent observability depends on host hooks and provider traffic reaching the local gateway. Relay cannot fully capture remote or cloud execution that bypasses the local host.
- Persistent Codex and Claude Code integrations use user-scoped
configuration and a shared loopback gateway. Use
nemo-relay runwhen a launch must retain project-specific configuration. - On Windows, a restrictive host Job Object can limit gateway reuse or prevent
persistent bootstrap. Codex can also make a cold-start
/modelsrequest before required MCP servers start; Relay retries the request. - Codex 0.143 does not expose
SessionEnd, and Codex multi-agent v2 encrypts delegated-task payloads that Relay cannot decrypt or reliably link. - The Node.js binding and package workflows require Node.js 24 or later.
- OpenClaw has public hook-backed telemetry. Its security and optimization coverage is partial because it does not own a managed execution path.
- The built-in
nemo_guardrailsplugin is deprecated and scheduled for removal in NeMo Relay 0.9. It remains available in 0.8: the remote backend inherits its configured service’s availability, latency, and policy behavior, and the local backend requires Python 3.11 or later andnemoguardrails==0.22.0. A replacement is not included in 0.8 and will target 0.9 or later. Removal will include the built-in component kind, the publicnemo_relay::plugins::nemo_guardrailsRust module, its CLI editor entry, and theguardrails-remoteCargo feature. - The PII redaction plugin currently supports its deterministic local backend; local-model backend configuration is reserved for future work.
- Pricing and optimization estimates depend on model names, token data, pricing sources, and freshness evidence. Missing or inconsistent evidence produces partial or absent cost fields rather than zero values.
- ATOF stream sinks and remote ATIF storage require reachable, correctly configured destinations. A failed stream sink does not stop file output or other active sinks.
- ATIF omits point-in-time marks. Use ATOF for the canonical mark stream;
the
gen_aiOpenTelemetry projection also omits marks. Thefullandopeninferenceprojections retain their fixed native mark handling. - Native dynamic plugins run in the Relay process and are not sandboxed. A
grpc-v1worker runs in a separate process, but that process is not a security sandbox. - Treat Python
LLMRequestobjects as immutable. Request middleware that changes content must return a new request object. - Native subscriber callbacks arrive asynchronously. Flush subscribers before depending on their side effects, captured events, files, or exporter output.
- OpenTelemetry endpoints use finite batch queues. A burst that fills an
endpoint queue drops completed spans without applying backpressure and can
leave an incomplete trace, including a missing root span. The SDK warns on
the first drop and reports the exact dropped-span count during graceful
shutdown. Configure
max_queue_size,max_export_batch_size, andscheduled_delay_millisindependently on each endpoint, or use the standardOTEL_BSP_*environment variables as process-wide fallbacks. A larger finite queue does not guarantee lossless telemetry. - Operational logging configuration and sink lifecycle are available, but broad operational log coverage across commands is not yet available.
Previous Releases
For previous release notes, release artifacts, and the complete PR-by-PR history, refer to GitHub Releases.