> This page is for version v1.2.0.
> For other versions, use one of these documentation indexes:
> - Latest (v1.5.1) (default): https://docs.nvidia.com/dynamo/latest/llms.txt
> - dev: https://docs.nvidia.com/dynamo/dev/llms.txt
> - v1.5.1: https://docs.nvidia.com/dynamo/v1.5.1/llms.txt
> - v1.5.0: https://docs.nvidia.com/dynamo/v1.5.0/llms.txt
> - v1.4.2: https://docs.nvidia.com/dynamo/v1.4.2/llms.txt
> - v1.4.1: https://docs.nvidia.com/dynamo/v1.4.1/llms.txt
> - v1.4.0: https://docs.nvidia.com/dynamo/v1.4.0/llms.txt
> - v1.3.0: https://docs.nvidia.com/dynamo/v1.3.0/llms.txt
> - v1.2.1: https://docs.nvidia.com/dynamo/v1.2.1/llms.txt
> - v1.2.0: https://docs.nvidia.com/dynamo/v1.2.0/llms.txt
> - v1.1.1: https://docs.nvidia.com/dynamo/v1.1.1/llms.txt
> - v1.1.0: https://docs.nvidia.com/dynamo/v1.1.0/llms.txt
> - v1.0.2: https://docs.nvidia.com/dynamo/v1.0.2/llms.txt
> - v1.0.1: https://docs.nvidia.com/dynamo/v1.0.1/llms.txt
> - v1.0.0: https://docs.nvidia.com/dynamo/v1.0.0/llms.txt
> - v0.9.1: https://docs.nvidia.com/dynamo/v-0-9-1/llms.txt
> - v0.9.0: https://docs.nvidia.com/dynamo/v-0-9-0/llms.txt
> - v0.8.1: https://docs.nvidia.com/dynamo/v-0-8-1/llms.txt
> - v0.8.0: https://docs.nvidia.com/dynamo/v-0-8-0/llms.txt
> - v0.7.1: https://docs.nvidia.com/dynamo/v-0-7-1/llms.txt
> - v0.7.0: https://docs.nvidia.com/dynamo/v-0-7-0/llms.txt

> For clean Markdown content of this page, append .md to this URL. For the complete documentation index, see https://docs.nvidia.com/dynamo/llms.txt. For full content including API reference and SDK examples, see https://docs.nvidia.com/dynamo/llms-full.txt.

# Quickstart

[![简体中文](https://fdr-prod-docs-files-public.s3.us-east-1.amazonaws.com/dynamo.docs.buildwithfern.com/2546607964ca3bb29badfbef8f4af73dee49dd6709242f791dfa79f7914b8f2e/pages-v1.2.0/assets/img/readme-zh-cn-link.svg?X-Amz-Algorithm=AWS4-HMAC-SHA256&X-Amz-Content-Sha256=UNSIGNED-PAYLOAD&X-Amz-Credential=AKIA6KXJSKKNFOCF7G4B%2F20261010%2Fus-east-1%2Fs3%2Faws4_request&X-Amz-Date=20261010T033640Z&X-Amz-Expires=604800&X-Amz-Signature=32e8ff41f47da86b965eb158dfb9570526839680b6422b741786110724cc3a2f&X-Amz-SignedHeaders=host&x-amz-checksum-mode=ENABLED&x-id=GetObject)](./quickstart.zh-CN.mdx)

## Choose Your Path

#### [Quickstart](/dynamo/dev/getting-started/quickstart)

You're here. Container fast path.

#### [Local Installation](/dynamo/dev/getting-started/local-installation)

Full walkthrough — PyPI, configuration.

#### [Kubernetes](/dynamo/dev/getting-started/kubernetes-deployment)

Kubernetes-native production path.

#### [Build from Source](/dynamo/dev/getting-started/building-from-source)

For contributors against `main`.

> **Note**
>
> Dynamo is backend-agnostic and Kubernetes-native without being Kubernetes-only. Use this container path to try the same frontend/router/worker stack locally; use the Kubernetes path when you want the operator, CRDs, Gateway API integration, autoscaling, scheduling, and cluster lifecycle management.

## Pull a Container

Containers have all dependencies pre-installed. Pick your backend:

#### SGLang

```bash
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.0.2
```

#### TensorRT-LLM

```bash
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.0.2
```

#### vLLM

```bash
docker run --gpus all --network host --rm -it nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.0.2
```

> **Warning**
>
> **Hugging Face token required for gated models.** Llama, Kimi, Qwen-VL, and other gated models require `HF_TOKEN` in your environment and accepting the model card's license on huggingface.co. Set `export HF_TOKEN=hf_…` before launching.

For container versions and tags, see [Release Artifacts](/dynamo/v1.2.0/resources/release-artifacts#container-images).

## Start the Frontend

In your container, start the OpenAI-compatible frontend on port 8000:

```bash
python3 -m dynamo.frontend --discovery-backend file
```

> **Tip**
>
> `--discovery-backend file` avoids needing etcd. To run frontend and worker in the same terminal, background each command with `> logfile.log 2>&1 &`.

## Start a Worker

In another terminal, launch a worker for your backend:

#### SGLang

```bash
python3 -m dynamo.sglang --model-path Qwen/Qwen3-0.6B --discovery-backend file
```

#### TensorRT-LLM

```bash
python3 -m dynamo.trtllm --model-path Qwen/Qwen3-0.6B --discovery-backend file
```

#### vLLM

```bash
python3 -m dynamo.vllm --model Qwen/Qwen3-0.6B --discovery-backend file \
  --kv-events-config '{"enable_kv_cache_events": false}'
```

## Verify and Test

Check the endpoint is up:

```bash
curl -sf http://localhost:8000/health && echo OK
```

If you see `OK`, send a chat completion:

**`Request`**

```bash title="Request"
curl localhost:8000/v1/chat/completions \
  -H "Content-Type: application/json" \
  -d '{"model": "Qwen/Qwen3-0.6B",
       "messages": [{"role": "user", "content": "Hello!"}],
       "max_tokens": 50}'
```

**`Response`**

```json title="Response"
{
  "id": "chatcmpl-...",
  "model": "Qwen/Qwen3-0.6B",
  "choices": [{
    "index": 0,
    "message": {"role": "assistant", "content": "Hello! How can I help you today?"},
    "finish_reason": "stop"
  }],
  "usage": {"prompt_tokens": 9, "completion_tokens": 10, "total_tokens": 19}
}
```

> **Info**
>
> Connection refused? The frontend takes a few seconds to start — retry. For production liveness and readiness probes, see [Health Checks](/dynamo/v1.2.0/user-guides/observability-local/health-checks).

## From the Digest

#### [Full-Stack Optimizations for Agentic Inference](/dynamo/dev/digest/agentic-inference)

How Dynamo optimizes for agentic workloads at three layers: the frontend API, the router, and KV cache management.

#### [Flash Indexer: Inter-Galactic KV Routing](/dynamo/dev/digest/flash-indexer)

How Dynamo's concurrent global index evolved through six iterations to sustain over 100M ops/sec.

## Dive Deeper

Pick a full install path from the [four options above](#choose-your-path), or explore how Dynamo works under the hood:

#### [Architecture](/dynamo/dev/design-docs/overall-architecture)

How the frontend, router, and workers fit together.

#### [Frontend Guide](/dynamo/dev/components/frontend)

Worker discovery, multi-model routing, OpenAI compat.

#### [KV Cache Aware Routing](/dynamo/dev/components/router/routing-concepts#kv-cache-routing)

How the router places requests for prefix reuse.

#### [Health Checks](/dynamo/dev/user-guides/observability-local/health-checks)

Liveness and readiness probes for production deployments.