Response Cache

View as Markdown

Use the response cache when the same LLM request is made more than once and the repeat should be served from a store instead of calling the provider again. An eligible request that matches an unexpired cache entry can be served from the store without calling the provider. Buffered hits preserve the stored response shape and usage fields; streaming hits replay an equivalent provider-native stream.

The cache is an optional response_cache section of the Adaptive plugin, not a standalone plugin kind. It is off until the section is present, applies to managed LLM calls without changing the execution API. By default, only requests with an explicit numeric temperature = 0 are eligible; set cache_nondeterministic = true to opt sampled requests into caching. Runtime backend errors fail open to a normal live call, while invalid configuration is rejected during validation.

namespace is required and defines one trusted cache-sharing domain. Do not use one namespace across mutually untrusted tenants or upstreams. header_allowlist does not replace this trust boundary.

When to Use It

  • Development and test loops that replay the same prompts — eligible repeats can reuse stored responses.
  • CI suites that exercise real models — cached repeats can avoid additional provider calls.
  • Demos and workshops — reuse stable stored answers while reducing provider calls.
  • A shared team cache — callers in one trusted sharing domain can use the same Redis store, namespace, and key prefix across processes and machines.
  • Gateway deployments — enable caching for an agent without touching its code.

Reuse requires the request to match exactly after normalization, so prompts that embed volatile content (timestamps, random IDs) will not repeat. The TTL is the maximum entry age, not guaranteed residency; a backend can evict or delete an entry sooner. Use bypass_rate to re-run a sample of cacheable calls live and refresh an entry when the new response is stored.

plugins.toml Example

1version = 1
2
3[[components]]
4kind = "adaptive"
5enabled = true
6
7[components.config]
8version = 1
9
10[components.config.response_cache]
11ttl_seconds = 3600 # maximum reuse age
12namespace = "dev-harness" # identifies one trusted cache-sharing domain
13bypass_rate = 0.0 # 0.1 would re-run 10% of cacheable calls live
14cache_nondeterministic = false # set true to cache nondeterministic requests
15
16[components.config.response_cache.backend]
17kind = "in_memory" # or "redis" for a shared cache

With this configuration the first eligible occurrence of a request runs the provider and attempts to store a complete answer. A later eligible, identical request within the TTL can be served from the store when that write succeeds and the entry remains present. The request’s provider surface is auto-detected from its shape, so there is nothing else to configure. For file discovery and precedence rules, refer to Plugin Configuration Files.

Plugin Configuration

Use plugin configuration when the application should let NeMo Relay own the cache lifecycle.

1import asyncio
2
3import nemo_relay
4
5adaptive_config = nemo_relay.adaptive.AdaptiveConfig(
6 response_cache=nemo_relay.adaptive.ResponseCacheConfig(
7 ttl_seconds=3600,
8 namespace="dev-harness",
9 ),
10)
11
12plugin_config = nemo_relay.plugin.PluginConfig(
13 components=[nemo_relay.adaptive.ComponentSpec(adaptive_config)]
14)
15
16report = nemo_relay.plugin.validate(plugin_config)
17if any(diagnostic["level"] == "error" for diagnostic in report["diagnostics"]):
18 raise RuntimeError(report["diagnostics"])
19
20def call_model(_request):
21 return {
22 "model": "gpt-4o",
23 "choices": [{
24 "message": {"role": "assistant", "content": "NeMo Relay instruments agent calls."},
25 "finish_reason": "stop",
26 }],
27 }
28
29async def main():
30 await nemo_relay.plugin.initialize(plugin_config)
31 try:
32 # Managed calls need no changes. The first call runs the provider;
33 # an identical repeat can be served from the cache.
34 request = nemo_relay.LLMRequest(
35 {},
36 {
37 "model": "gpt-4o",
38 "messages": [{"role": "user", "content": "What is NeMo Relay?"}],
39 "temperature": 0,
40 },
41 )
42 await nemo_relay.llm.execute("openai", request, call_model)
43 await nemo_relay.llm.execute("openai", request, call_model)
44 finally:
45 await nemo_relay.plugin.clear_async()
46
47asyncio.run(main())

Canonical plugin documents and plugins.toml use snake_case. Python response-cache fields follow that convention, while Node.js uses camelCase at its response-cache helper surface. ComponentSpec and AdaptiveRuntime serialize those fields to canonical keys. Python uses BackendSpec.redis(url, key_prefix=...); Node.js uses adaptive.redisBackend(url, keyPrefix). Both Redis helpers default the prefix to "nemo_relay:", overriding the value in the fields table. Configure the same key_prefix in every process and binding that should share entries.

Manual API

Use the manual runtime API when an integration needs to own the adaptive lifecycle directly instead of activating the top-level plugin component.

1import asyncio
2
3import nemo_relay
4
5adaptive_config = nemo_relay.adaptive.AdaptiveConfig(
6 response_cache=nemo_relay.adaptive.ResponseCacheConfig(namespace="dev-harness"),
7)
8
9async def main():
10 runtime = nemo_relay.adaptive.AdaptiveRuntime(adaptive_config.to_dict())
11 await runtime.register()
12 try:
13 # Run instrumented application work here.
14 runtime.wait_for_idle()
15 finally:
16 await runtime.shutdown()
17
18asyncio.run(main())

What Gets Cached

Only complete, replayable answers are stored:

  • A response with a non-null error or a status such as failed, cancelled, incomplete, or in_progress is never stored.
  • Stateful requests bypass the cache entirely: a request that opts into server-side persistence (a non-null store value other than false, or an OpenAI Responses call without an explicit store = false), continues a stored interaction (previous_response_id), or references server-side state (conversation, container). Their answers depend on state the cache key cannot see.
  • Streaming calls are cached too. On a miss the live chunks are forwarded to the consumer while being assembled into one aggregate response — the same shape a buffered call stores, so buffered and streaming calls share one keyspace. On a hit the stored answer is replayed as provider-native chunks (OpenAI Chat deltas, OpenAI Responses lifecycle events, or Anthropic Messages events), so strict streaming clients parse it like a live stream. Provider-terminal token-limited Chat (finish_reason = "length") and Anthropic (stop_reason = "max_tokens") answers can be stored. OpenAI Responses answers with status = "incomplete", streams without a terminal event, and streams whose content cannot be replayed faithfully (for example thinking blocks) are never stored. A streaming request whose surface cannot be inferred runs live, uncached.

Streaming publication is write-behind: Relay can report end-of-stream before the backend write finishes. wait_for_idle() does not wait for cache writes. An immediate repeat can therefore run live, and concurrent identical misses can each call the provider because the cache does not coalesce them.

A buffered hit returns the stored response unchanged. A streaming hit replays an equivalent provider-native stream. Available saved-token and estimated-cost values are reported on the response_cache mark, never by editing a buffered response body.

Cache Keys

Two requests hit the same entry when they are the same request after normalization, under the default key_strategy = "exact_request":

  • The request is decoded to its normalized form and fingerprinted with SHA-256, so provider-shaped differences that mean the same thing collapse to one key. When a request cannot be decoded faithfully, the raw body is fingerprinted instead — that fallback can only cost a miss, never a wrong hit.
  • Field order and whitespace never matter (RFC 8785 canonicalization).
  • Only the built-in noise fields stream, user, metadata, and store are dropped. All other normalized or raw provider controls remain in the key, including service_tier; there is no configurable skip list.
  • OpenAI Chat requests preserve whether the caller used max_tokens or max_completion_tokens, because providers can treat the two fields differently.
  • Tool-call IDs are normalized, so randomly generated per-call IDs do not fragment the keyspace.
  • Request headers stay out of the key unless named in header_allowlist. Allowlist every trusted, non-secret response-affecting header; an omitted header does not partition the key. Known auth headers are rejected, but validation cannot recognize every custom credential name, so never allowlist credentials.
  • The provider name and required namespace partition every key. A Switchyard-selected backend ID is also partitioned automatically. Without Switchyard, a provider name cannot distinguish upstreams that reuse that name, so use separate configurations and namespaces for different trusted upstream domains.
  • Requests containing integers outside the exactly representable RFC 8785 range (less than -2^53 or greater than 2^53) bypass the cache.

Observability

Every cache decision emits a response_cache mark with data.status set to one of:

StatusMeaning
hitServed the stored answer; the provider was skipped.
missNo entry was served; the call ran live. After an ordinary lookup miss, Relay attempts to store a cacheable result.
bypassThe request is not cacheable, or the bypass_rate sampler chose to run live.

Mark attributes use nemo_relay.response_cache.*: backend, surface, key_hash (the sha256:… fingerprint), ttl_ms, and age_ms as applicable; saved_tokens and saved_cost_usd appear on hits when they can be derived. A reason appears on bypasses and store-error misses (for example sampled, stateful_store, store_error, or stream_no_codec). Cache marks never include prompts, answers, or credentials.

nemo-relay doctor reports the cache state: not configured when the section is absent, configured but disabled (adaptive plugin disabled) when the adaptive component is off, on; backend '<kind>' reachable when healthy, and a failure when the config is invalid or the backend is unreachable.

Fields

FieldDefaultNotes
ttl_seconds3600Maximum entry age in seconds; a backend can evict an entry sooner.
namespace"" (unconfigured)Required non-empty trust-domain partition folded into every key. Empty or whitespace-only values are rejected.
priority50LLM execution intercept priority. Lower values run earlier.
bypass_rate0.0Probability in [0.0, 1.0] of running a cacheable call live. A sampled call attempts to refresh the entry when its result is cacheable and the write succeeds.
cache_nondeterministicfalseOnly requests with an explicit numeric temperature = 0 are eligible. Set true to cache and reuse sampled responses.
key_strategy"exact_request"The only supported strategy: reuse requires the same normalized request.
header_allowlist[]Trusted, non-secret response-affecting headers folded into the key (case-insensitive). Known auth headers are rejected.
backend.kind"in_memory""in_memory", or "redis" (requires building with the redis-backend feature).
backend.config.max_bytes256 MiBIn-memory size budget; the oldest entries are evicted first.
backend.config.urlRedis connection URL. Required for the redis backend.
backend.config.key_prefix"nemo-relay:llm-cache:"Prefix for keys in Redis.

For a gateway that uses Switchyard, switchyard.priority must be lower than response_cache.priority. To derive keys before ACG rewrites requests, set response_cache.priority lower than acg.priority. With all three components, priorities of 0, 40, and 50, respectively, satisfy both orderings.

Cached responses are stored unredacted. PII sanitize guardrails rewrite emitted telemetry, never payloads, so the store holds full response bodies. Cache entries can also store provider and model diagnostics plus the key fingerprint; they do not store full request bodies or headers. A shared Redis backend must be trusted and access-controlled. Use a separate configuration and namespace for each mutually untrusted tenant or upstream domain.

Common Validation Failures

  • namespace is empty or whitespace-only, ttl_seconds is 0, or bypass_rate is outside [0.0, 1.0].
  • key_strategy is not "exact_request".
  • header_allowlist names an auth header such as authorization or x-api-key.
  • in_memory max_bytes is zero or is not an integer.
  • backend.kind is unknown; or redis has a missing, non-string, or whitespace-only backend.config.url, uses a non-string key_prefix, or is unavailable because Relay was built without the redis-backend feature.
  • Gateway Switchyard priority is equal to or greater than response_cache.priority.