This page is for version v0.9.0.
For other versions, use one of these documentation indexes:
- Latest (v1.5.1) (default): https://docs.nvidia.com/dynamo/latest/llms.txt
- dev: https://docs.nvidia.com/dynamo/dev/llms.txt
- v1.5.1: https://docs.nvidia.com/dynamo/v1.5.1/llms.txt
- v1.5.0: https://docs.nvidia.com/dynamo/v1.5.0/llms.txt
- v1.4.2: https://docs.nvidia.com/dynamo/v1.4.2/llms.txt
- v1.4.1: https://docs.nvidia.com/dynamo/v1.4.1/llms.txt
- v1.4.0: https://docs.nvidia.com/dynamo/v1.4.0/llms.txt
- v1.3.0: https://docs.nvidia.com/dynamo/v1.3.0/llms.txt
- v1.2.1: https://docs.nvidia.com/dynamo/v1.2.1/llms.txt
- v1.2.0: https://docs.nvidia.com/dynamo/v1.2.0/llms.txt
- v1.1.1: https://docs.nvidia.com/dynamo/v1.1.1/llms.txt
- v1.1.0: https://docs.nvidia.com/dynamo/v1.1.0/llms.txt
- v1.0.2: https://docs.nvidia.com/dynamo/v1.0.2/llms.txt
- v1.0.1: https://docs.nvidia.com/dynamo/v1.0.1/llms.txt
- v1.0.0: https://docs.nvidia.com/dynamo/v1.0.0/llms.txt
- v0.9.1: https://docs.nvidia.com/dynamo/v-0-9-1/llms.txt
- v0.9.0: https://docs.nvidia.com/dynamo/v-0-9-0/llms.txt
- v0.8.1: https://docs.nvidia.com/dynamo/v-0-8-1/llms.txt
- v0.8.0: https://docs.nvidia.com/dynamo/v-0-8-0/llms.txt
- v0.7.1: https://docs.nvidia.com/dynamo/v-0-7-1/llms.txt
- v0.7.0: https://docs.nvidia.com/dynamo/v-0-7-0/llms.txt
Enable SGLang Hierarchical Cache (HiCache)
This guide shows how to enable SGLang’s Hierarchical Cache (HiCache) inside Dynamo.
1) Start the SGLang worker with HiCache enabled
—enable-hierarchical-cache : Enables hierarchical KV cache/offload
—hicache-ratio : The ratio of the size of host KV cache memory pool to the size of device pool. Lower this number if your machine has less CPU memory.
—hicache-write-policy : Write policy (e.g., write_through for synchronous host writes)
—hicache-storage-backend : Host storage backend for HiCache (e.g., nixl). NIXL selects the concrete store automatically; see PR #8488
Then, start the frontend:
2) Send a single request
3) (Optional) Benchmarking
Run the perf script: