SGLang

SGLang engines run in Dynamo’s distributed runtime with disaggregated serving, KV-aware routing, and request cancellation.

View as Markdown

Use the Latest Release

We recommend using the latest stable release of Dynamo to avoid breaking changes.


Dynamo SGLang integrates SGLang engines into Dynamo’s distributed runtime, enabling disaggregated serving, KV-aware routing, and request cancellation while maintaining full compatibility with SGLang’s native engine arguments. It supports LLM inference, embedding models, multimodal vision models, and diffusion-based generation (LLM, image, video).

Prerequisites

  • CUDA toolkit headers for bare-metal builds (e.g. nvcc, cuda_runtime.h). See CUDA Requirements. Not required when running the pre-built sglang-runtime container.

  • HF_TOKEN for gated models. Export it on every node that pulls the model weights, and accept the model license on the Hugging Face model page before launch:

    $export HF_TOKEN=hf_...

Installation

Install Latest Release

We recommend using uv to install:

$uv venv --python 3.12 --seed
$uv pip install --prerelease=allow "ai-dynamo[sglang]"

This installs the latest stable release of Dynamo with the compatible SGLang version.

Install for Development

Docker

Two paths are supported. Pick the one that matches how you plan to develop.

Pull and launch the published sglang-runtime image from NGC. See release artifacts for the current tag and CUDA variants.

$docker run --gpus all -it --rm \
> --network host --shm-size=10G \
> --ulimit memlock=-1 --ulimit stack=67108864 \
> --ulimit nofile=65536:65536 \
> --cap-add CAP_SYS_PTRACE --ipc host \
> -v $HOME/.cache/huggingface:/home/dynamo/.cache/huggingface \
> nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.2.1

Mount the host Hugging Face cache (-v $HOME/.cache/huggingface:/home/dynamo/.cache/huggingface) so each container restart doesn’t re-download model weights. The container runs as user dynamo (UID 1000), which is why the in-container path is /home/dynamo/.cache/huggingface.

Build from source inside upstream SGLang container

Pull and launch the upstream SGLang image, then build Dynamo from source inside it:

$docker run --gpus all -it --rm \
> --network host --shm-size=10G \
> --ulimit memlock=-1 --ulimit stack=67108864 \
> --ulimit nofile=65536:65536 \
> --ipc host \
> lmsysorg/sglang:v{sglang_version}

Install build dependencies and Rust inside the container:

$apt-get update -qq && apt-get install -y -qq \
> build-essential libclang-dev curl git > /dev/null 2>&1
$
$curl --proto '=https' --tlsv1.2 -sSf https://sh.rustup.rs | sh -s -- -y
$source "$HOME/.cargo/env"
$
$pip install maturin[patchelf]

Clone and build Dynamo:

$cd /sgl-workspace/
$git clone https://github.com/ai-dynamo/dynamo.git
$cd dynamo
$
$cd lib/bindings/python/
$maturin build -o /tmp
$pip install /tmp/ai_dynamo_runtime*.whl
$
$cd /sgl-workspace/dynamo/
$pip install -e .

Feature Support Matrix

FeatureStatusNotes
Disaggregated ServingPrefill/decode separation with NIXL KV transfer
KV-Aware Routing
SLA-Based Planner
Multimodal SupportImage via EPD, E/PD, E/P/D patterns
Diffusion ModelsLLM diffusion, image, and video generation
Request CancellationAggregated full; disaggregated decode-only
Graceful ShutdownDiscovery unregister + grace period
ObservabilityMetrics, tracing, and Grafana dashboards

Feature Interactions

SGLang is optimized for high-throughput serving with fast primitives, providing robust support for disaggregated serving, KV-aware routing, and request migration. The matrix below shows which feature pairs are validated to work together.

Legend: ✅ Supported  |  🚧 Work in Progress / Experimental / Limited

FeatureDisaggregated ServingKV-Aware RoutingSLA-Based PlannerKV Block ManagerMultimodalRequest MigrationRequest CancellationLoRATool CallingSpeculative Decoding
Disaggregated Serving
KV-Aware Routing
SLA-Based Planner
KV Block Manager🚧🚧🚧
Multimodal21🚧
Request Migration🚧
Request Cancellation🚧3🚧🚧
LoRA🚧
Tool Calling🚧
Speculative Decoding🚧🚧🚧🚧🚧

Notes:

  1. Multimodal + KV-Aware Routing: Not supported. (Source)
  2. Multimodal Patterns: Supports simple Aggregated EPD, E/PD, and E/P/D patterns. Traditional Disagg EP/D is not supported. (Source)
  3. Request Cancellation: Cancellation during the remote prefill phase is not supported in disaggregated mode. (Source)
  4. Speculative Decoding: Code hooks exist (spec_decode_stats in publisher), but no examples or documentation yet.

Quick Start

Python / CLI Deployment

Start infrastructure services for local development:

$docker compose -f dev/docker-compose.yml up -d

Launch an aggregated serving deployment:

$cd $DYNAMO_HOME/examples/backends/sglang
$./launch/agg.sh

Verify the deployment:

$curl localhost:8000/v1/chat/completions \
> -H "Content-Type: application/json" \
> -d '{
> "model": "Qwen/Qwen3-0.6B",
> "messages": [{"role": "user", "content": "Explain why Roger Federer is considered one of the greatest tennis players of all time"}],
> "stream": true,
> "max_tokens": 30
> }'

Disaggregated Serving

Launch a disaggregated Qwen3-0.6B deployment (smallest model, useful for plumbing validation):

$cd $DYNAMO_HOME/examples/backends/sglang
$./launch/disagg.sh

Performance caveat: Qwen3-0.6B is small enough that the disaggregated pathway is dominated by transport overhead and will often look slower than aggregated. Use it for plumbing validation, not benchmarks. Switch to Qwen3-32B-FP8 or larger for realistic disagg numbers.

Multi-Node TP

SGLang supports multi-node tensor parallelism via the native --dist-init-addr, --nnodes, and --node-rank flags. See SGLang server arguments for the canonical reference; the same flags work with python -m dynamo.sglang. For a Kubernetes deployment example, see disagg-multinode.yaml.

Kubernetes Deployment

You can deploy SGLang with Dynamo on Kubernetes using a DynamoGraphDeployment. For more details, see the SGLang Kubernetes Deployment Guide.

Next Steps