> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/relay/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/relay/_mcp/server.

# gRPC Worker Protocol Overview

> Understand the local `grpc-v1` contract and canonical protocol definition for NeMo Relay worker plugins.

`grpc-v1` is the stable out-of-process plugin protocol for
`plugin.kind = "worker"`. Python, Rust, and other local executables can
implement the `nemo.relay.worker.v1` API. Relay starts every worker and connects
to it through local endpoints. Remote worker endpoints are not supported.

Use the Rust or Python SDK unless you need another runtime. Refer to [gRPC Worker Plugin Concepts](/build-plugins/dynamic-plugins/grpc-worker/about) for runtime
choices and lifecycle guidance.

## Service Contract

Workers implement the `PluginWorker` service:

* `Handshake` and `Health` identify a ready worker.
* `Validate` returns configuration diagnostics.
* `Register` returns declarative subscriber, guardrail, and intercept
  registrations.
* `Invoke` and `InvokeStream` run registered behavior.
* `CancelInvocation` requests cancellation, and `Shutdown` requests process
  termination.

Relay implements `RelayHostRuntime` for worker-initiated operations. It lets a
worker emit marks, manage scopes and isolated scope stacks, and call tool, LLM,
or LLM-stream continuations during execution intercepts.

This page summarizes the stable contract. Refer to the [canonical protocol definition](https://github.com/NVIDIA/NeMo-Relay/blob/0.5.0-alpha.20260706/crates/worker-proto/proto/nemo/relay/worker/v1/plugin_worker.proto)
for every message and RPC field required to implement another runtime.

## PluginWorker RPCs

The worker implements the following RPCs:

| RPC                | Request and response                     | Purpose                                                                                        |
| ------------------ | ---------------------------------------- | ---------------------------------------------------------------------------------------------- |
| `Handshake`        | `HandshakeRequest` → `HandshakeResponse` | Confirms plugin identity, `grpc-v1`, SDK/runtime details, and supported registration surfaces. |
| `Health`           | `HealthRequest` → `HealthResponse`       | Reports whether the worker is ready.                                                           |
| `Validate`         | `ValidateRequest` → `ValidateResponse`   | Validates the component `config` envelope and returns diagnostics or `WorkerError`.            |
| `Register`         | `RegisterRequest` → `RegisterResponse`   | Returns declarative registrations or `WorkerError`.                                            |
| `Invoke`           | `InvokeRequest` → `InvokeResponse`       | Runs a subscriber, guardrail, or non-streaming intercept.                                      |
| `InvokeStream`     | `InvokeRequest` → `stream StreamChunk`   | Runs a streaming LLM intercept.                                                                |
| `CancelInvocation` | `CancelInvocationRequest` → `WorkerAck`  | Requests cooperative cancellation of an invocation.                                            |
| `Shutdown`         | `ShutdownRequest` → `WorkerAck`          | Requests worker shutdown after Relay removes its proxy callbacks.                              |

Every request carries the activation ID and authentication token. `Validate`
and `Register` also carry the plugin ID and component configuration.

## Registration and Invocation

On success, `RegisterResponse` contains `Registration` records with a local
name, surface, priority, and `break_chain` value. A failed registration can
return `WorkerError` without registrations. The supported surfaces are:

* `SUBSCRIBER`
* `TOOL_SANITIZE_REQUEST_GUARDRAIL`, `TOOL_SANITIZE_RESPONSE_GUARDRAIL`,
  `TOOL_CONDITIONAL_EXECUTION_GUARDRAIL`, `TOOL_REQUEST_INTERCEPT`, and
  `TOOL_EXECUTION_INTERCEPT`
* `LLM_SANITIZE_REQUEST_GUARDRAIL`, `LLM_SANITIZE_RESPONSE_GUARDRAIL`,
  `LLM_CONDITIONAL_EXECUTION_GUARDRAIL`, `LLM_REQUEST_INTERCEPT`,
  `LLM_EXECUTION_INTERCEPT`, and `LLM_STREAM_EXECUTION_INTERCEPT`

`InvokeRequest` identifies the registration, surface, invocation, optional
continuation, and scope context. Its payload is one of an event, tool
invocation, or LLM invocation. `InvokeResponse` returns an empty result, JSON
result, guardrail result, LLM request-intercept result, tool-execution result,
or `WorkerError`. `InvokeStream` emits JSON chunks or `WorkerError` chunks.

Every LLM sanitizer invocation includes a directional context with tagged codec
identity: `none`, `builtin(id)`, `runtime(id)`, or `opaque`. Worker SDKs expose
`context.resolve_codec()` for active codecs. The resulting invocation-scoped
proxy supports request decode/encode or response decode and calls Relay through
the host runtime; the capability identifier is protocol-internal and expires
when the callback completes. `resolve_codec()` returns no proxy only when no
codec is active. Runtime and opaque codecs remain resolvable.

Request and response sanitizer handlers always receive `(payload, context)`.
Return the sanitized payload to continue the chain, or return no payload to
omit the observability payload and annotation without changing the
client-visible value. Rust worker sanitizer callbacks are async for every
sanitizer surface: mark, scope, tool, and LLM. Python worker sanitizers can
return either an immediate value or an awaitable.

#### Rust

```rust
ctx.register_llm_sanitize_request_guardrail(
    "normalize-request",
    10,
    |request, context| async move {
        let Some(codec) = context.resolve_codec() else {
            return Ok(Some(request));
        };
        let mut annotated = codec.decode(&request).await?;
        annotated.messages.clear();
        Ok(Some(codec.encode(&annotated, &request).await?))
    },
);
```

#### Python

```python
async def normalize_request(request, context):
    codec = context.resolve_codec()
    if codec is None:
        return request
    annotated = await codec.decode(request)
    annotated["messages"] = []
    return await codec.encode(annotated, request)

ctx.register_llm_sanitize_request_guardrail(
    "normalize-request",
    normalize_request,
    priority=10,
)
```

## RelayHostRuntime RPCs

Relay implements these RPCs for worker callbacks:

| RPC group               | RPCs                                                                       |
| ----------------------- | -------------------------------------------------------------------------- |
| Marks and scopes        | `EmitMark`, `PushScope`, `PopScope`                                        |
| Isolated scope stacks   | `CreateScopeStack`, `DropScopeStack`                                       |
| Execution continuations | `ToolNext`, `LlmNext`, `LlmStreamNext`                                     |
| Codec capabilities      | `DecodeLlmCodecRequest`, `EncodeLlmCodecRequest`, `DecodeLlmCodecResponse` |

Every host-runtime request also carries the activation ID and authentication
token. Scope operations include a `ScopeContext`; continuation calls include
the continuation ID that Relay supplied for the active intercept. Codec
operations additionally require the unforgeable capability ID supplied for the
current sanitizer invocation. Relay rejects missing, forged, expired,
wrong-direction, or activation-mismatched capabilities. Codec transformation
failures are non-retryable worker errors.

A worker may call an execution continuation repeatedly or concurrently while
the corresponding `Invoke` or `InvokeStream` callback is active. Each call gets
an isolated scope-stack branch containing the scopes visible at invocation.
The worker must finish continuation calls before returning its middleware
result; Relay removes the continuation ID and cancels unfinished calls when the
callback settles. For `InvokeStream`, the returned worker stream extends the
active callback lifetime until it closes, so it can call `LlmStreamNext` lazily.
A stream successfully returned by `LlmStreamNext` keeps its normal streaming
lifetime.

## Authentication and Endpoints

Relay supplies an activation ID, an activation token, the worker endpoint, and
the host endpoint when it starts the process. SDKs attach the activation ID and
token to protocol calls. The host rejects requests with an invalid activation ID
or token.

On Unix platforms, Relay uses local Unix sockets. On other platforms, Relay
uses loopback TCP endpoints. Workers must not accept arbitrary remote endpoint
configuration.

## Payloads

Relay data values use this envelope:

```proto
message JsonEnvelope {
  string schema = 1;
  bytes json = 2;
}
```

Use `JsonEnvelope` for Relay DTOs. Protobuf defines protocol control flow and
does not duplicate Relay event, tool, LLM, scope, or diagnostic models.

## Activation and Shutdown

1. Run `nemo-relay plugins validate <plugin-id>` before enabling or running a
   plugin to check its manifest, trust evidence, and optional static
   configuration schema.
2. Relay creates local endpoints, starts the worker, and completes health and
   handshake checks.
3. Relay sends component configuration to `Validate`. Error diagnostics stop
   initialization after the worker process starts.
4. Relay obtains registrations from `Register` and installs proxy callbacks.
   Relay rolls back installed callbacks if initialization fails.
5. Relay removes proxy callbacks before it sends `Shutdown` to the worker.

## Errors

gRPC status communicates transport, authentication, malformed protocol
requests, and some stream failures. Registration and unary callback failures
can return structured `WorkerError` values. A stream callback failure can
terminate the gRPC stream with a status error. Worker implementations must
handle both error forms.