Response Cache
Use the response cache when the same LLM request is made more than once and the repeat should be served from a store instead of calling the provider again. An eligible request that matches an unexpired cache entry can be served from the store without calling the provider. Buffered hits preserve the stored response shape and usage fields; streaming hits replay an equivalent provider-native stream.
The cache is an optional response_cache section of the
Adaptive plugin, not a standalone plugin
kind. It is off until the section is present, applies to
managed LLM calls without
changing the execution API. By default, only requests with an explicit numeric
temperature = 0 are eligible; set cache_nondeterministic = true to opt
sampled requests into caching. Runtime backend errors fail open to a normal
live call, while invalid configuration is rejected during validation.
namespace is required and defines one trusted cache-sharing domain. Do not
use one namespace across mutually untrusted tenants or upstreams.
header_allowlist does not replace this trust boundary.
When to Use It
- Development and test loops that replay the same prompts — eligible repeats can reuse stored responses.
- CI suites that exercise real models — cached repeats can avoid additional provider calls.
- Demos and workshops — reuse stable stored answers while reducing provider calls.
- A shared team cache — callers in one trusted sharing domain can use the same Redis store, namespace, and key prefix across processes and machines.
- Gateway deployments — enable caching for an agent without touching its code.
Reuse requires the request to match exactly after normalization, so prompts
that embed volatile content (timestamps, random IDs) will not repeat. The TTL
is the maximum entry age, not guaranteed residency; a backend can evict or
delete an entry sooner. Use bypass_rate to re-run a sample of cacheable calls
live and refresh an entry when the new response is stored.
plugins.toml Example
With this configuration the first eligible occurrence of a request runs the provider and attempts to store a complete answer. A later eligible, identical request within the TTL can be served from the store when that write succeeds and the entry remains present. The request’s provider surface is auto-detected from its shape, so there is nothing else to configure. For file discovery and precedence rules, refer to Plugin Configuration Files.
Plugin Configuration
Use plugin configuration when the application should let NeMo Relay own the cache lifecycle.
Python
Node.js
Rust
Canonical plugin documents and plugins.toml use snake_case. Python
response-cache fields follow that convention, while Node.js uses camelCase at
its response-cache helper surface. ComponentSpec and AdaptiveRuntime
serialize those fields to canonical keys. Python uses
BackendSpec.redis(url, key_prefix=...); Node.js uses
adaptive.redisBackend(url, keyPrefix). Both Redis helpers default the prefix
to "nemo_relay:", overriding the value in the fields table. Configure the
same key_prefix in every process and binding that should share entries.
Manual API
Use the manual runtime API when an integration needs to own the adaptive lifecycle directly instead of activating the top-level plugin component.
Python
Node.js
Rust
What Gets Cached
Only complete, replayable answers are stored:
- A response with a non-null
erroror astatussuch asfailed,cancelled,incomplete, orin_progressis never stored. - Stateful requests bypass the cache entirely: a request that opts into
server-side persistence (a non-null
storevalue other thanfalse, or an OpenAI Responses call without an explicitstore = false), continues a stored interaction (previous_response_id), or references server-side state (conversation,container). Their answers depend on state the cache key cannot see. - Streaming calls are cached too. On a miss the live chunks are forwarded to
the consumer while being assembled into one aggregate response — the same
shape a buffered call stores, so buffered and streaming calls share one
keyspace. On a hit the stored answer is replayed as provider-native chunks
(OpenAI Chat deltas, OpenAI Responses lifecycle events, or Anthropic
Messages events), so strict streaming clients parse it like a live stream.
Provider-terminal token-limited Chat (
finish_reason = "length") and Anthropic (stop_reason = "max_tokens") answers can be stored. OpenAI Responses answers withstatus = "incomplete", streams without a terminal event, and streams whose content cannot be replayed faithfully (for example thinking blocks) are never stored. A streaming request whose surface cannot be inferred runs live, uncached.
Streaming publication is write-behind: Relay can report end-of-stream before
the backend write finishes. wait_for_idle() does not wait for cache writes.
An immediate repeat can therefore run live, and concurrent identical misses
can each call the provider because the cache does not coalesce them.
A buffered hit returns the stored response unchanged. A streaming hit replays
an equivalent provider-native stream. Available saved-token and estimated-cost
values are reported on the response_cache mark, never by editing a buffered
response body.
Cache Keys
Two requests hit the same entry when they are the same request after
normalization, under the default key_strategy = "exact_request":
- The request is decoded to its normalized form and fingerprinted with SHA-256, so provider-shaped differences that mean the same thing collapse to one key. When a request cannot be decoded faithfully, the raw body is fingerprinted instead — that fallback can only cost a miss, never a wrong hit.
- Field order and whitespace never matter (RFC 8785 canonicalization).
- Only the built-in noise fields
stream,user,metadata, andstoreare dropped. All other normalized or raw provider controls remain in the key, includingservice_tier; there is no configurable skip list. - OpenAI Chat requests preserve whether the caller used
max_tokensormax_completion_tokens, because providers can treat the two fields differently. - Tool-call IDs are normalized, so randomly generated per-call IDs do not fragment the keyspace.
- Request headers stay out of the key unless named in
header_allowlist. Allowlist every trusted, non-secret response-affecting header; an omitted header does not partition the key. Known auth headers are rejected, but validation cannot recognize every custom credential name, so never allowlist credentials. - The provider name and required namespace partition every key. A Switchyard-selected backend ID is also partitioned automatically. Without Switchyard, a provider name cannot distinguish upstreams that reuse that name, so use separate configurations and namespaces for different trusted upstream domains.
- Requests containing integers outside the exactly representable RFC 8785
range (less than
-2^53or greater than2^53) bypass the cache.
Observability
Every cache decision emits a response_cache mark with
data.status set to one of:
Mark attributes use nemo_relay.response_cache.*: backend, surface,
key_hash (the sha256:… fingerprint), ttl_ms, and age_ms as applicable;
saved_tokens and saved_cost_usd appear on hits when they can be derived. A
reason appears on bypasses and store-error misses (for example sampled,
stateful_store, store_error, or stream_no_codec). Cache marks never
include prompts, answers, or credentials.
nemo-relay doctor reports the cache state: not configured when the section
is absent, configured but disabled (adaptive plugin disabled) when the
adaptive component is off, on; backend '<kind>' reachable when healthy, and
a failure when the config is invalid or the backend is unreachable.
Fields
For a gateway that uses Switchyard, switchyard.priority must be lower than
response_cache.priority. To derive keys before ACG rewrites requests, set
response_cache.priority lower than acg.priority. With all three components,
priorities of 0, 40, and 50, respectively, satisfy both orderings.
Cached responses are stored unredacted. PII sanitize guardrails rewrite emitted telemetry, never payloads, so the store holds full response bodies. Cache entries can also store provider and model diagnostics plus the key fingerprint; they do not store full request bodies or headers. A shared Redis backend must be trusted and access-controlled. Use a separate configuration and namespace for each mutually untrusted tenant or upstream domain.
Common Validation Failures
namespaceis empty or whitespace-only,ttl_secondsis0, orbypass_rateis outside[0.0, 1.0].key_strategyis not"exact_request".header_allowlistnames an auth header such asauthorizationorx-api-key.in_memorymax_bytesis zero or is not an integer.backend.kindis unknown; orredishas a missing, non-string, or whitespace-onlybackend.config.url, uses a non-stringkey_prefix, or is unavailable because Relay was built without theredis-backendfeature.- Gateway Switchyard priority is equal to or greater than
response_cache.priority.