Releases (machine-readable)

Generated release, compatibility, and artifact data for agents and automation
View as Markdown

This page is a plain-markdown rendering of components/releases.data.ts, the single source of truth behind Compatibility, Release Artifacts, Model Early Access Builds, and the Release Notes timeline. It is regenerated by scripts/gen_llms_tables.py at every release bump — do not edit the tables below by hand. Append .md to this page’s URL for a clean markdown export. The same data ships in-repo as JSON (docs/fern/assets/releases.json) and as an Atom feed (docs/fern/assets/releases-atom.xml).

Current stable release: v1.3.0 (Jul 20, 2026; container tag 1.3.0, wheel version 1.3.0.post1).

Releases

VersionKindDateSGLangTensorRT-LLMvLLMNIXL (SGL / TRT / vLLM)UCXNotesDelta
main (ToT)development head-0.5.151.3.0rc210.25.11.3.0 / 1.0.1 / 1.1.0---
v1.3.0stableJul 20, 20260.5.141.3.0rc190.23.01.3.0 / 1.0.1 / 1.1.01.20.xrelease notesCUDA 12 container images discontinued; EFA variants go multi-arch as -efa; GA wheels published as 1.3.0.post1 (containers stay :1.3.0); UCX 1.20.x.
v1.3.0-dev.1platform-previewJun 9, 20260.5.12.post11.3.0rc170.22.01.0.1 / 0.10.1 / 1.1.0-release notesFull-platform preview of v1.3.0: complete runtime matrix, wheels on pypi.nvidia.com, crates, and Helm charts. Superseded by v1.3.0 GA.
v1.2.1patchJun 13, 20260.5.111.3.0rc140.20.11.0.1 / 0.10.1 / 0.10.1-release notesPatch release. Same backend pins as v1.2.0.
v1.2.0stableJun 2, 20260.5.111.3.0rc140.20.11.0.1 / 0.10.1 / 0.10.11.20.0release notes603 PRs from 82 authors. DGD/DGDR promoted to v1beta1; CRTC default approximate KV router; inter-pod GMS sidecar; Dynamo Snapshot on CRI-O / OpenShift; UCX 1.20.0.
v1.2.0-deepseek-v4-dev.3model-buildMay 9, 2026upstream DSv4 preview-0.20.1- / - / 0.10.1-release notesDeepSeek-V4 Blackwell preview; vLLM + SGLang containers only.
v1.2.0-deepseek-v4-dev.2model-buildMay 1, 2026upstream DSv4 preview-0.20.0- / - / 0.10.1-release notesDeepSeek-V4 Blackwell preview; vLLM + SGLang containers only.
v1.1.1patchMay 5, 20260.5.10.post11.3.0rc110.19.01.0.1 / 0.10.1 / 0.10.1-release notesPatch release. Same backend pins as v1.1.0.
v1.1.0stableMay 1, 20260.5.10.post11.3.0rc110.19.01.0.1 / 0.10.1 / 0.10.11.20release notesPlanner split into its own dynamo-planner image (artifact boundary change). First 1.y.z publication of dynamo-protocols on crates.io; dynamo-async-openai deprecated at final 1.0.2.
v1.1.0-dev.3platform-previewApr 18, 20260.5.10.post11.3.0rc110.19.01.0.1 / 0.10.1 / 0.10.1-release notesPartial platform preview: TRT-LLM runtime image + wheels only.
v1.1.0-dev.2platform-previewApr 9, 20260.5.91.3.0rc90.19.01.0.1 / 0.10.1 / 0.10.1-release notesPartial platform preview: SGLang + TRT-LLM runtime images + wheels.
v1.1.0-dev.1platform-previewMar 17, 20260.5.91.3.0rc5.post10.17.11.0.1 / 0.10.1 / 0.10.1-release notesPlatform preview: runtime matrix, wheels on pypi.nvidia.com, Helm charts.
v1.0.2patchApr 22, 20260.5.91.3.0rc5.post10.16.00.10.1 / 0.10.1 / 0.10.1-release notesNo artifact additions or removals versus v1.0.0.
v1.0.1patchMar 16, 20260.5.91.3.0rc5.post10.16.00.10.1 / 0.10.1 / 0.10.1-release notesNo artifact additions or removals versus v1.0.0.
v1.0.0stableMar 12, 20260.5.91.3.0rc5.post10.16.00.10.1 / 0.10.1 / 0.10.1-release notessnapshot-agent image and EFA variants for vLLM and TRT-LLM (AMD64 only). First publish of dynamo-mocker and dynamo-kv-router crates. snapshot Helm chart added (preview); deprecated dynamo-crds dropped from the publish stream.
v0.9.1patchMar 4, 20260.5.81.3.0rc30.14.10.9.0 / 0.9.0 / 0.9.0-release notesNo artifact additions or removals versus v0.9.0.
v0.9.0stableFeb 11, 20260.5.81.3.0rc10.14.10.9.0 / 0.9.0 / 0.9.0-release notesFirst publish of dynamo-tokens crate. Deprecated dynamo-graph Helm chart dropped from the publish stream.
v0.8.1.post3patch-0.5.6.post21.2.0rc6.post30.12.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post3. Same CUDA support as v0.8.1.
v0.8.1.post2patch-0.5.6.post21.2.0rc6.post20.12.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post2. Same CUDA support as v0.8.1.
v0.8.1.post1patch-0.5.6.post21.2.0rc6.post10.12.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post1. Same CUDA support as v0.8.1.
v0.8.1patchJan 23, 20260.5.6.post21.2.0rc6.post10.12.00.8.0 / 0.8.0 / 0.8.0-release notesPost trains .post1/.post2/.post3 republished the TRT-LLM runtime image and PyPI wheels only; each carried a distinct TRT-LLM pin (see the v0.8.1.post1/.post2/.post3 rows).
v0.8.0stableJan 15, 20260.5.6.post21.2.0rc6.post10.12.00.8.0 / 0.8.0 / 0.8.0-release notesdynamo-frontend image and CUDA 13 variants for vLLM and SGLang. First publish of dynamo-memory and dynamo-config crates.
v0.7.1patchDec 15, 20250.5.4.post31.2.0rc30.11.00.8.0 / 0.8.0 / 0.8.0-release notes-
v0.7.0.post1patch-0.5.4.post31.2.0rc30.11.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.7.0: TensorRT-LLM pin advanced to 1.2.0rc3 (v0.7.0 shipped 1.2.0rc2). Same CUDA support as v0.7.0.
v0.7.0stableNov 26, 20250.5.4.post31.2.0rc20.11.00.8.0 / 0.8.0 / 0.8.0-release notes-
v0.6.1.post1patch-0.5.3.post21.1.0rc50.11.00.6.0 / 0.6.0 / 0.6.0--Post-train of v0.6.1: same backend pins as v0.6.1. Same CUDA support as v0.6.1.
v0.6.1patchNov 6, 20250.5.3.post21.1.0rc50.11.00.6.0 / 0.6.0 / 0.6.0-release notes-
v0.6.0stableOct 28, 20250.5.3.post21.1.0rc50.11.00.6.0 / 0.6.0 / 0.6.0-release notesOldest release tracked on this page.

Release highlights (stable releases):

  • v1.3.0: Tool-calling and reasoning overhaul, RL rollout serving, the largest Router buildout to date, SLA-driven Planner autoscaling, and production GPU Memory Service on Kubernetes.
  • v1.2.0: DGD/DGDR v1beta1, CRTC as the default KV router, inter-pod GPU Memory Service, Dynamo Snapshot on CRI-O/OpenShift, and DeepSeek-V4 recipes on vLLM.
  • v1.1.0: Resilient KV routing at scale, Anthropic Messages API support, performance modeling and offline replay, and the multimodal embedding cache.
  • v1.0.0: First GA release: unified configuration, Kubernetes production readiness, multimodal serving, and the agents surface.

CUDA toolkit and minimum driver history

DynamoBackendCUDA ToolkitMin DriverNote
1.3.0SGLang13.0580.xx+-
1.3.0TensorRT-LLM13.1580.xx+-
1.3.0vLLM13.0580.xx+-
1.2.1SGLang12.9575.xx+-
1.2.1SGLang13.0580.xx+-
1.2.1TensorRT-LLM13.1580.xx+-
1.2.1vLLM12.9575.xx+-
1.2.1vLLM13.0580.xx+-
1.2.0SGLang12.9575.xx+-
1.2.0SGLang13.0580.xx+-
1.2.0TensorRT-LLM13.1580.xx+-
1.2.0vLLM12.9575.xx+-
1.2.0vLLM13.0580.xx+-
1.1.1SGLang12.9575.xx+-
1.1.1SGLang13.0580.xx+-
1.1.1TensorRT-LLM13.1580.xx+-
1.1.1vLLM12.9575.xx+-
1.1.1vLLM13.0580.xx+-
1.1.0SGLang12.9575.xx+-
1.1.0SGLang13.0580.xx+-
1.1.0TensorRT-LLM13.1580.xx+-
1.1.0vLLM12.9575.xx+-
1.1.0vLLM13.0580.xx+-
1.0.2SGLang12.9575.xx+-
1.0.2SGLang13.0580.xx+-
1.0.2TensorRT-LLM13.1580.xx+-
1.0.2vLLM12.9575.xx+-
1.0.2vLLM13.0580.xx+-
1.0.1SGLang12.9575.xx+-
1.0.1SGLang13.0580.xx+-
1.0.1TensorRT-LLM13.1580.xx+-
1.0.1vLLM12.9575.xx+-
1.0.1vLLM13.0580.xx+-
1.0.0SGLang12.9575.xx+-
1.0.0SGLang13.0580.xx+-
1.0.0TensorRT-LLM13.1580.xx+-
1.0.0vLLM12.9575.xx+-
1.0.0vLLM13.0580.xx+-
0.9.1SGLang12.9575.xx+-
0.9.1TensorRT-LLM13.0580.xx+-
0.9.1vLLM12.9575.xx+-
0.9.0SGLang12.9575.xx+-
0.9.0TensorRT-LLM13.0580.xx+-
0.9.0vLLM12.9575.xx+-
0.8.1SGLang12.9575.xx+-
0.8.1SGLang13.0580.xx+Experimental
0.8.1TensorRT-LLM13.0580.xx+-
0.8.1vLLM12.9575.xx+-
0.8.1vLLM13.0580.xx+Experimental
0.8.0SGLang12.9575.xx+-
0.8.0SGLang13.0580.xx+Experimental
0.8.0TensorRT-LLM13.0580.xx+-
0.8.0vLLM12.9575.xx+-
0.8.0vLLM13.0580.xx+Experimental
0.7.1SGLang12.8570.xx+-
0.7.1TensorRT-LLM13.0580.xx+-
0.7.1vLLM12.9575.xx+-
0.7.0SGLang12.9575.xx+-
0.7.0TensorRT-LLM13.0580.xx+-
0.7.0vLLM12.8570.xx+-
  • Patch versions (e.g. v0.8.1.post1, v0.7.0.post1) have the same CUDA support as their base version.
  • Early access v1.1.0-dev.* images follow the same CUDA matrix as v1.0.2. The v1.2.0-deepseek-v4-dev.3 vLLM container is CUDA 13.0 multi-arch; the SGLang containers split by arch (CUDA 12.9 on amd64, CUDA 13.0 on arm64).
  • Experimental CUDA 13 images are not published for all versions.

Feature support by backend (v1.3.0)

FeatureSGLangTensorRT-LLMvLLM
Disaggregated ServingSupportedSupportedSupported (Prefill/decode separation with NIXL KV transfer)
KV-Aware RoutingSupportedSupportedSupported
SLA-Based PlannerSupportedSupportedSupported
KV Block ManagerExperimental (Work in progress across all combinations)SupportedSupported
Multimodal (Image)Supported (Not compatible with KV-aware routing. Disagg patterns: EPD, E/PD, E/P/D (not traditional EP/D))Supported (Image URLs + pre-computed embeddings. Disagg: EP/D + E/P/D. KV-aware routing via dedicated MM Router Worker (requires KV event publishing))Supported (With KV-aware routing, image-aware routing on documented paths)
Multimodal (Video)SupportedNot supportedSupported (Video input with frame sampling)
Multimodal (Audio)Not supportedNot supportedExperimental (Qwen2-Audio, experimental)
Request MigrationSupportedSupported (Work in progress with multimodal)Supported
Request CancellationExperimental (Remote-prefill-phase cancellation not supported in disaggregated mode)Supported with caveat (Engine temporarily not notified of cancellations — resources for cancelled requests are not freed (known issue))Supported
LoRANot supportedNot supportedSupported (Dynamic load/unload; KV-aware routing supports adapter affinity)
Tool CallingSupportedSupportedSupported
Speculative DecodingExperimental (Code hooks exist; no examples or docs yet)SupportedSupported (Eagle3)
Dynamo SnapshotSupportedNot supportedSupported

Artifact inventory (v1.3.0)

CategoryNameDescriptionMetaTags / install
containervllm-runtimevLLM backend runtimevLLM v0.23.0 · CUDA 13.0 · AMD64/ARM64nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.3.0; nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.3.0-efa
containersglang-runtimeSGLang backend runtimeSGLang v0.5.14 · CUDA 13.0 · AMD64/ARM64nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.3.0; nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.3.0-efa
containertensorrtllm-runtimeTensorRT-LLM backend runtimeTRT-LLM v1.3.0rc19 · CUDA 13.1 · AMD64/ARM64nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.3.0; nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.3.0-efa
containerdynamo-frontendOpenAI-compatible API gateway with Endpoint Prediction Protocol (EPP)AMD64/ARM64nvcr.io/nvidia/ai-dynamo/dynamo-frontend:1.3.0
containerdynamo-plannerStandalone Planner used by Profiler jobs and Planner podsAMD64/ARM64nvcr.io/nvidia/ai-dynamo/dynamo-planner:1.3.0
containerkubernetes-operatorOperator that manages Dynamo deployments and CRDsAMD64/ARM64nvcr.io/nvidia/ai-dynamo/kubernetes-operator:1.3.0
containersnapshot-agent (Preview)Fast GPU worker recovery via CRIUAMD64/ARM64nvcr.io/nvidia/ai-dynamo/snapshot-agent:1.3.0
wheelai-dynamoMain package with backend integrations (vLLM, SGLang, TRT-LLM)Python 3.10–3.12 · Linux (glibc v2.28+)uv pip install ai-dynamo==1.3.0.post1
wheelai-dynamo-runtimeCore Python bindings for the Dynamo runtimePython 3.10–3.12 · Linux (glibc v2.28+)uv pip install ai-dynamo-runtime==1.3.0.post1
wheelkvbmKV Block Manager for disaggregated KV cachePython 3.10–3.12 · Linux (glibc v2.28+)uv pip install kvbm==1.3.0.post1
helmdynamo-platformPlatform services (etcd, NATS) and the Dynamo Operator for a Dynamo cluster-helm install dynamo-platform oci://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/dynamo-platform --version 1.3.0
helmsnapshotSnapshot DaemonSet for fast GPU worker recovery-helm install snapshot oci://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/snapshot --version 1.3.0
cratedynamo-runtimeCore distributed runtime libraryMSRV Rust v1.82cargo add dynamo-runtime@1.3.0
cratedynamo-llmLLM inference engineMSRV Rust v1.82cargo add dynamo-llm@1.3.0
cratedynamo-protocolsAsync OpenAI-compatible API clientMSRV Rust v1.82cargo add dynamo-protocols@1.3.0
cratedynamo-async-openai (Deprecated)Legacy OpenAI client; use dynamo-protocolsMSRV Rust v1.82 · final releasecargo add dynamo-async-openai@1.0.2
cratedynamo-parsersProtocol parsers (SSE, JSON streaming)MSRV Rust v1.82cargo add dynamo-parsers@1.3.0
cratedynamo-memoryMemory management utilitiesMSRV Rust v1.82cargo add dynamo-memory@1.3.0
cratedynamo-configConfiguration managementMSRV Rust v1.82cargo add dynamo-config@1.3.0
cratedynamo-tokensTokenizer bindings for LLM inferenceMSRV Rust v1.82cargo add dynamo-tokens@1.3.0
cratedynamo-tokenizersTokenizer library for LLM inferenceMSRV Rust v1.82cargo add dynamo-tokenizers@1.3.0
cratedynamo-mockerInference engine simulator for benchmarkingMSRV Rust v1.82cargo add dynamo-mocker@1.3.0
cratedynamo-kv-routerKV-aware request routing libraryMSRV Rust v1.82cargo add dynamo-kv-router@1.3.0
cratekvbm-logicalLogical layer for the KV Block ManagerMSRV Rust v1.82cargo add kvbm-logical@1.3.0

Known artifact issues

ReleaseArtifactIssueStatus
v0.9.0dynamo-platform-0.9.0Helm chart sets operator image to 0.7.1 instead of 0.9.0.Fixed in v0.9.0.post1
v0.8.1vllm-runtime:0.8.1-cuda13Container fails to launch.Known issue
v0.8.1sglang-runtime:0.8.1-cuda13, vllm-runtime:0.8.1-cuda13Multimodality not expected to work on ARM64. Works on AMD64.Known limitation
v0.8.0sglang-runtime:0.8.0-cuda13CuDNN installation issue caused PyTorch v2.9.1 compatibility problems with nn.Conv3d — performance degradation and excessive memory usage in multimodal workloads.Fixed in v0.8.1 (#5461)

Crates: first published version on crates.io

CrateFirst versionDate
dynamo-runtime0.1.02025-03-18
dynamo-llm0.2.02025-05-01
dynamo-async-openai0.4.12025-08-27
dynamo-parsers0.5.02025-09-18
dynamo-memory0.8.02026-01-15
dynamo-config0.8.02026-01-15
dynamo-tokens0.9.02026-02-12
dynamo-mocker1.0.02026-03-13
dynamo-kv-router1.0.02026-03-13
dynamo-protocols1.1.02026-05-04
dynamo-tokenizers1.2.02026-06-02

Model early-access builds

ModelTagRelease lineRuntimesShippedGA pathStatusCoverage (images / wheels / helm / crates)
Inkling1.4.0-inkling-dev.1v1.4.0sglang-runtimeJul 17, 2026Dev-only · v1.4.0 lineFirst build on the v1.4.0 line; targets the next stable release.yes / no / no / no
GLM-5.21.3.0-glm-5.2-dev.1v1.3.0sglang-runtimeJul 20, 2026Dev-onlyContainer carries SGLang cherry-picks (stability, config parsing, model support) opened upstream but not yet in a released SGLang.yes / no / no / no
MiniMax-M31.3.0-minimax-m3-dev.1v1.3.0vllm-runtime, sglang-runtime, tensorrtllm-runtimeJun 12, 2026Promoted → :1.3.0Dynamo changes and the M2 tool-calling fix are in release/1.3.0; the recipes run on the stock :1.3.0 containers.yes / no / no / no
DeepSeek-V41.3.0-deepseek-v4-dev.1v1.3.0tensorrtllm-runtimeJun 6, 2026Recipe in v1.3.0DeepSeek-V4 Flash and Pro recipes ship in v1.3.0 on the standard TensorRT-LLM release container.yes / no / no / no
Nemotron-3-Ultra1.3.0-nemotron-ultra-dev.1v1.3.0vllm-runtimeJun 5, 2026Dev-onlyFour un-upstreamed vLLM patches; requires pinned flags VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernel and —no-enable-flashinfer-autotune.yes / no / no / no
Nemotron-3-Super1.3.0-nemotron-super-dev.1v1.3.0vllm-runtimeJun 4, 2026Promoted → :1.3.0Both container patches are in the vLLM v0.23.0 that v1.3.0 ships; the recipe runs on the stock vllm-runtime:1.3.0.yes / no / no / no
Kimi-K2.61.3.0-kimi-k2.6-dev.1v1.3.0vllm-runtimeJun 4, 2026Promoted → :1.3.0The build’s only container patch is in vLLM v0.23.0; the recipes run on the stock vllm-runtime:1.3.0.yes / no / no / no
Cosmos-31.3.0-cosmos3-dev.1v1.3.0vllm-runtimeJun 1, 2026Dev-onlyDynamo #10132 (Cosmos3 support in the vLLM-Omni backend) is open, not merged — v1.3.0 containers cannot run Cosmos3.yes / no / no / no
DeepSeek-V4 preview1.2.0-deepseek-v4-dev.3v1.2.0vllm-runtime, sglang-runtimeMay 9, 2026Superseded — recipe in v1.3.0Blackwell (B200 + GB200) preview; per-arch/CUDA tags (e.g. vllm-runtime:1.2.0-deepseek-v4-cuda13-dev.3). Superseded by the v1.3.0 recipe.yes / no / no / no
DeepSeek-V4 preview1.2.0-deepseek-v4-dev.2v1.2.0vllm-runtime, sglang-runtimeMay 1, 2026Superseded — recipe in v1.3.0Blackwell preview on vLLM v0.20.0 (native DSv4 support); superseded by dev.3.yes / no / no / no
DeepSeek-V4 preview1.2.0-sglang-deepseek-v4-dev.1v1.2.0sglang-runtimeApr 25, 2026Superseded — recipe in v1.3.0Earliest DSv4 preview (SGLang, B200 only); superseded by dev.2/dev.3.yes / no / no / no

Platform-preview artifact coverage

PreviewImagesWheelsHelmCrates
v1.3.0-dev.1yesyesyesyes
v1.1.0-dev.3yesyesnono
v1.1.0-dev.2yesyesnono
v1.1.0-dev.1yesyesyesno

Platform support

  • GPU architectures: Blackwell, Hopper, Ada Lovelace, Ampere
  • OS: Ubuntu 24.04 (x86_64, ARM64) — Supported
  • OS: Ubuntu 22.04 (x86_64) — Supported
  • OS: CentOS Stream 9 (x86_64) — Experimental
  • CSP: AWS — Amazon Linux 2023 (x86_64) — Supported
  • CPU architectures: x86_64, ARM64 (Ubuntu 24.04 only)
  • Wheels: Wheels are built in a manylinux_2_28-compatible environment and validated on CentOS Stream 9 and Ubuntu 22.04/24.04. Other Linux distributions are expected to work but are not officially verified.

Release statistics

ReleasePRsContributorsFirst-time contributorsBreaking changesKnown issues
v1.3.0930125232410
v1.2.060382-511
v1.1.089611312820
v1.0.0-90344114

Nightlies

ai-dynamo and ai-dynamo-runtime nightly builds from main publish wheels tagged *.devYYYYMMDD (since Apr 24, 2026). Install with pip or uv using —pre and the NVIDIA extra-index pattern shown above.