Offload KV Cache Locally

Extend KV cache capacity beyond GPU memory when running Dynamo from a container or Python environment
View as Markdown

KV cache offloading moves reusable KV blocks from GPU memory to host memory or storage. This tutorial starts an aggregated Dynamo deployment with one offloading backend, sends requests through it, and checks that the backend is running.

1

Prepare Your Environment

Complete the local installation and start Dynamo’s NATS and etcd services before continuing. Run the commands from the root of the cloned Dynamo repository.

If docker run opened a shell in a prebuilt runtime container, you are already in the correct environment and can skip package installation when the table below marks the backend as included. For a container running in the background, open a shell with docker exec; the repository is available at /workspace:

docker exec -it <container-name> bash
cd /workspace

If you installed Dynamo into a Python environment, activate that environment before continuing.

2

Choose a Backend

Select one offloading backend for each worker. LMCache, FlexKV, and HiCache are alternatives, not cache layers to stack together. KV-aware routing and NIXL-based prefill/decode transfer are separate features that can operate alongside the selected backend.

BackendEngineCache tiersPrebuilt Dynamo runtime containerPython environment
LMCachevLLMHost and persistent L2 adaptersInherited from the upstream vLLM image; verify compatibility for the image tag and platformInstall lmcache
HiCacheSGLangHost and optional external storageIncluded with SGLangIncluded with ai-dynamo[sglang]
FlexKVvLLM in this tutorialHost, SSD, cloud storageRequires a custom image or a source build inside the containerBuild FlexKV from source

The Dynamo Operator does not install these packages. Local containers and Kubernetes deployments get them from the selected runtime image. The Python environment path installs packages on the host instead.

3

Start Offloading

Choose a backend and follow its tab. The primary examples use Qwen/Qwen3-0.6B and one GPU; the optional disaggregated FlexKV example uses two GPUs. The vLLM launch scripts start the Dynamo frontend and workers, and stop all processes when you press Ctrl+C.

The current Dynamo vLLM runtime is based on an upstream vLLM image that provides LMCache on supported platforms. Confirm that the selected image contains a compatible build:

python -c "import lmcache; print(lmcache.__version__)"

For a compatible x86_64 Python environment, install LMCache with:

uv pip install lmcache

LMCache publishes x86_64 wheels built for specific CUDA versions. For Arm64 or a mismatched PyTorch/CUDA stack, build LMCache from source by following the LMCache installation guide.

Some Dynamo vLLM image tags include an LMCache build that is incompatible with LMCacheMPConnector. If startup fails with RuntimeError: Unsupported GPUKVFormat, build LMCache from a version containing LMCache pull request #3282.

Start the out-of-process LMCache server and an aggregated vLLM worker:

cd examples/backends/vllm
LMCACHE_L1_SIZE_GB=16 ./launch/agg_lmcache_mp.sh

The script starts lmcache server, waits for its health endpoint, and connects the vLLM worker with LMCacheMPConnector. Inspect its metrics from another terminal:

curl -s localhost:8080/metrics | grep '^lmcache_mp_'

For server flags and persistent L2 adapters, see the LMCache MP configuration reference.

4

Send a Request

After the frontend listens on port 8000, send a request from another terminal:

curl localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "Qwen/Qwen3-0.6B",
"messages": [{"role": "user", "content": "Explain how a KV cache reduces repeated computation."}],
"stream": false,
"max_tokens": 30
}'

Send the request again to create an opportunity to reuse its prompt prefix. A short request confirms the deployment path but does not establish a performance improvement. Use a representative workload with long, repeated prefixes to measure Time To First Token (TTFT) and cache hit rate.

Next Steps

  • Add KV-aware routing when multiple workers can reuse cached prefixes.
  • Combine one offloading backend with disaggregated serving. In that topology, the offloading connector manages storage while NIXL transfers blocks from prefill to decode workers.
  • To make a custom inference engine visible to the KV router, implement KV event publishing.