> This page is for version v1.2.0.
> For other versions, use one of these documentation indexes:
> - Latest (v1.5.1) (default): https://docs.nvidia.com/dynamo/latest/llms.txt
> - dev: https://docs.nvidia.com/dynamo/dev/llms.txt
> - v1.5.1: https://docs.nvidia.com/dynamo/v1.5.1/llms.txt
> - v1.5.0: https://docs.nvidia.com/dynamo/v1.5.0/llms.txt
> - v1.4.2: https://docs.nvidia.com/dynamo/v1.4.2/llms.txt
> - v1.4.1: https://docs.nvidia.com/dynamo/v1.4.1/llms.txt
> - v1.4.0: https://docs.nvidia.com/dynamo/v1.4.0/llms.txt
> - v1.3.0: https://docs.nvidia.com/dynamo/v1.3.0/llms.txt
> - v1.2.1: https://docs.nvidia.com/dynamo/v1.2.1/llms.txt
> - v1.2.0: https://docs.nvidia.com/dynamo/v1.2.0/llms.txt
> - v1.1.1: https://docs.nvidia.com/dynamo/v1.1.1/llms.txt
> - v1.1.0: https://docs.nvidia.com/dynamo/v1.1.0/llms.txt
> - v1.0.2: https://docs.nvidia.com/dynamo/v1.0.2/llms.txt
> - v1.0.1: https://docs.nvidia.com/dynamo/v1.0.1/llms.txt
> - v1.0.0: https://docs.nvidia.com/dynamo/v1.0.0/llms.txt
> - v0.9.1: https://docs.nvidia.com/dynamo/v-0-9-1/llms.txt
> - v0.9.0: https://docs.nvidia.com/dynamo/v-0-9-0/llms.txt
> - v0.8.1: https://docs.nvidia.com/dynamo/v-0-8-1/llms.txt
> - v0.8.0: https://docs.nvidia.com/dynamo/v-0-8-0/llms.txt
> - v0.7.1: https://docs.nvidia.com/dynamo/v-0-7-1/llms.txt
> - v0.7.0: https://docs.nvidia.com/dynamo/v-0-7-0/llms.txt

> For clean Markdown content of this page, append .md to this URL. For the complete documentation index, see https://docs.nvidia.com/dynamo/llms.txt. For full content including API reference and SDK examples, see https://docs.nvidia.com/dynamo/llms-full.txt.

# Feature Guides

Use these guides after you have Dynamo running and want to improve serving behavior, operate a deployment, or adapt Dynamo to a new workload.

## Recommended path

Most deployments start with the core performance loop:

| Step | Guide | Use when |
|---|---|---|
| 1 | [KV Cache Aware Routing](/dynamo/v1.2.0/user-guides/kv-cache-aware-routing) | Route requests to workers that already hold useful KV cache. |
| 2 | [Disaggregated Serving](/dynamo/v1.2.0/user-guides/disaggregated-serving) | Scale prefill and decode workers independently. |
| 3 | [KV Cache Offloading](/dynamo/v1.2.0/user-guides/kv-cache-offloading) | Extend usable cache capacity beyond GPU memory. |
| 4 | [Benchmarking](/dynamo/v1.2.0/user-guides/benchmarking) | Compare configurations before you move to production. |

## Where to go next

| Goal | Start with |
|---|---|
| Make serving more resilient | [Fault Tolerance](/dynamo/v1.2.0/user-guides/fault-tolerance) |
| Monitor local deployments | [Observability (Local)](/dynamo/v1.2.0/user-guides/observability-local) |
| Reproduce traffic without a full engine | [Mocker Engine Simulation](../mocker/mocker.md) |
| Add structured model outputs | [Tool Calling](/dynamo/v1.2.0/user-guides/parsing/tool-call-parsing-dynamo) and [Reasoning](/dynamo/v1.2.0/user-guides/parsing/reasoning-parsing-dynamo) |
| Build agent workloads | [Agents](/dynamo/v1.2.0/user-guides/agents) |
| Serve specialized workloads | [LoRA Adapters](/dynamo/v1.2.0/user-guides/lo-ra-adapters), [Multimodal](/dynamo/v1.2.0/user-guides/multimodal), and [Diffusion](/dynamo/v1.2.0/user-guides/diffusion) |

For cluster deployments, pair these guides with the [Kubernetes Deployment](/dynamo/v1.2.0/kubernetes-deployment/start-here/kubernetes-quickstart) docs. The same features can be explored locally, then expressed through Dynamo's Kubernetes-native CRDs and operator when you move to a shared GPU cluster.