Releases (machine-readable)
Releases (machine-readable)
This page is a plain-markdown rendering of components/releases.data.ts, the single source of truth behind Compatibility, Release Artifacts, Model Early Access Builds, and the Release Notes timeline. It is regenerated by scripts/gen_llms_tables.py at every release bump — do not edit the tables below by hand. Append .md to this page’s URL for a clean markdown export. Fetch the same data as JSON or an Atom feed from the docs-website branch, which the publish pipeline regenerates on every docs publish; the files are no longer committed to main.
Current stable release: v1.4.2 (Aug 28, 2026; container tag 1.4.2, wheel version 1.4.2).
Releases
| Version | Kind | Date | SGLang | TensorRT-LLM | vLLM | NIXL (SGL / TRT / vLLM) | UCX | Notes | Delta |
|---|---|---|---|---|---|---|---|---|---|
| main (ToT) | development head | - | 0.5.18 | 1.3.0rc25 | 0.28.0 | 1.4.0 / 1.3.1 / 1.3.2 | - | - | - |
| v1.4.2 | patch | Aug 28, 2026 | 0.5.16 | 1.3.0rc22 | 0.26.0 | 1.3.0 / 1.3.1 / 1.3.2 | 1.21.x | release notes | Patch release and the first Dynamo Enterprise release: a curated set of release artifacts publishes under the -enterprise suffix on NGC, eligible for NVIDIA Enterprise Support, with no functional or binary differences from the open-source artifacts. Fixes NIXL loader-path resolution in the Frontend and SGLang Runtime images, removes the unused Nsight EFA metrics plugin, and tightens dependency pins (pillow v12.3.0 floor, plotext below v6, EFA Installer v1.50). Backend pins are unchanged from v1.4.0. |
| v1.4.1 | patch | Aug 21, 2026 | 0.5.16 | 1.3.0rc22 | 0.26.0 | 1.3.0 / 1.3.1 / 1.3.2 | 1.21.x | release notes | Patch release. Adds the classify and pooling endpoints, forwards logprob_token_ids through the OpenAI frontend, reconciles request-path overload marks in the Router, and fixes NIXL writable buffers for vLLM. All three Go modules move to Go 1.26.6 with aligned x/net and grpc. Backend pins are unchanged from v1.4.0. |
| v1.4.0 | stable | Aug 14, 2026 | 0.5.16 | 1.3.0rc22 | 0.26.0 | 1.3.0 / 1.3.1 / 1.3.2 | 1.21.x | release notes | Audit subsystem migrated into request trace (DYN_AUDIT_* honored as legacy aliases); HTTP header capture in trace records is an explicit fail-closed allowlist; deprecated multimodal worker flags and vLLM worker-role flags removed; runtime images no longer bundle software video decoders (H.264/H.265 decodes via NVDEC); UCX 1.21.x. |
| v1.3.1 | patch | Aug 5, 2026 | 0.5.14 | 1.3.0rc19 | 0.23.0 | 1.3.2 / 1.0.1 / 1.1.0 | 1.20.x | release notes | Patch release. Fixes disaggregated SGLang serving over AWS EFA on GB200: the SGLang EFA runtime moves to NIXL 1.3.2 and all three EFA images to EFA Installer 1.49.0. Backend pins are otherwise unchanged from v1.3.0. |
| v1.3.0 | stable | Jul 20, 2026 | 0.5.14 | 1.3.0rc19 | 0.23.0 | 1.3.0 / 1.0.1 / 1.1.0 | 1.20.x | release notes | CUDA 12 container images discontinued; EFA variants retagged from -efa-amd64 to -efa (the images were already multi-arch — the old suffix was misleading); GA wheels published as 1.3.0.post1 (containers stay :1.3.0); UCX 1.20.x. |
| v1.3.0-dev.1 | platform-preview | Jun 9, 2026 | 0.5.12.post1 | 1.3.0rc17 | 0.22.0 | 1.0.1 / 0.10.1 / 1.1.0 | - | release notes | Full-platform preview of v1.3.0: complete runtime matrix, wheels on pypi.nvidia.com, crates, and Helm charts. Superseded by v1.3.0 GA. |
| v1.2.1 | patch | Jun 13, 2026 | 0.5.11 | 1.3.0rc14 | 0.20.1 | 1.0.1 / 0.10.1 / 0.10.1 | - | release notes | Patch release. Same backend pins as v1.2.0. |
| v1.2.0 | stable | Jun 2, 2026 | 0.5.11 | 1.3.0rc14 | 0.20.1 | 1.0.1 / 0.10.1 / 0.10.1 | 1.20.0 | release notes | 603 PRs from 82 authors. DGD/DGDR promoted to v1beta1; CRTC default approximate KV router; inter-pod GMS sidecar; Dynamo Snapshot on CRI-O / OpenShift; UCX 1.20.0. |
| v1.2.0-deepseek-v4-dev.3 | model-build | May 9, 2026 | upstream DSv4 preview | - | 0.20.1 | - / - / 0.10.1 | - | release notes | DeepSeek-V4 Blackwell preview; vLLM + SGLang containers only. |
| v1.2.0-deepseek-v4-dev.2 | model-build | May 1, 2026 | upstream DSv4 preview | - | 0.20.0 | - / - / 0.10.1 | - | release notes | DeepSeek-V4 Blackwell preview; vLLM + SGLang containers only. |
| v1.1.1 | patch | May 5, 2026 | 0.5.10.post1 | 1.3.0rc11 | 0.19.0 | 1.0.1 / 0.10.1 / 0.10.1 | - | release notes | Patch release. Same backend pins as v1.1.0. |
| v1.1.0 | stable | May 1, 2026 | 0.5.10.post1 | 1.3.0rc11 | 0.19.0 | 1.0.1 / 0.10.1 / 0.10.1 | 1.20 | release notes | Planner split into its own dynamo-planner image (artifact boundary change). First 1.y.z publication of dynamo-protocols on crates.io; dynamo-async-openai deprecated at final 1.0.2. |
| v1.1.0-dev.3 | platform-preview | Apr 18, 2026 | 0.5.10.post1 | 1.3.0rc11 | 0.19.0 | 1.0.1 / 0.10.1 / 0.10.1 | - | release notes | Partial platform preview: TRT-LLM runtime image + wheels only. |
| v1.1.0-dev.2 | platform-preview | Apr 9, 2026 | 0.5.9 | 1.3.0rc9 | 0.19.0 | 1.0.1 / 0.10.1 / 0.10.1 | - | release notes | Partial platform preview: SGLang + TRT-LLM runtime images + wheels. |
| v1.1.0-dev.1 | platform-preview | Mar 17, 2026 | 0.5.9 | 1.3.0rc5.post1 | 0.17.1 | 1.0.1 / 0.10.1 / 0.10.1 | - | release notes | Platform preview: runtime matrix, wheels on pypi.nvidia.com, Helm charts. |
| v1.0.2 | patch | Apr 22, 2026 | 0.5.9 | 1.3.0rc5.post1 | 0.16.0 | 0.10.1 / 0.10.1 / 0.10.1 | - | release notes | No artifact additions or removals versus v1.0.0. |
| v1.0.1 | patch | Mar 16, 2026 | 0.5.9 | 1.3.0rc5.post1 | 0.16.0 | 0.10.1 / 0.10.1 / 0.10.1 | - | release notes | No artifact additions or removals versus v1.0.0. |
| v1.0.0 | stable | Mar 12, 2026 | 0.5.9 | 1.3.0rc5.post1 | 0.16.0 | 0.10.1 / 0.10.1 / 0.10.1 | - | release notes | snapshot-agent image and EFA variants for vLLM and TensorRT-LLM. First publish of dynamo-mocker and dynamo-kv-router crates. snapshot Helm chart added (preview); deprecated dynamo-crds dropped from the publish stream. |
| v0.9.1 | patch | Mar 4, 2026 | 0.5.8 | 1.3.0rc3 | 0.14.1 | 0.9.0 / 0.9.0 / 0.9.0 | - | release notes | No artifact additions or removals versus v0.9.0. |
| v0.9.0 | stable | Feb 11, 2026 | 0.5.8 | 1.3.0rc1 | 0.14.1 | 0.9.0 / 0.9.0 / 0.9.0 | - | release notes | First publish of dynamo-tokens crate. Deprecated dynamo-graph Helm chart dropped from the publish stream. |
| v0.8.1.post3 | patch | - | 0.5.6.post2 | 1.2.0rc6.post3 | 0.12.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | - | Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post3. Same CUDA support as v0.8.1. |
| v0.8.1.post2 | patch | - | 0.5.6.post2 | 1.2.0rc6.post2 | 0.12.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | - | Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post2. Same CUDA support as v0.8.1. |
| v0.8.1.post1 | patch | - | 0.5.6.post2 | 1.2.0rc6.post1 | 0.12.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | - | Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post1. Same CUDA support as v0.8.1. |
| v0.8.1 | patch | Jan 23, 2026 | 0.5.6.post2 | 1.2.0rc6.post1 | 0.12.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | release notes | Post trains .post1/.post2/.post3 republished the TRT-LLM runtime image and PyPI wheels only; each carried a distinct TRT-LLM pin (see the v0.8.1.post1/.post2/.post3 rows). |
| v0.8.0 | stable | Jan 15, 2026 | 0.5.6.post2 | 1.2.0rc6.post1 | 0.12.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | release notes | dynamo-frontend image and CUDA 13 variants for vLLM and SGLang. First publish of dynamo-memory and dynamo-config crates. |
| v0.7.1 | patch | Dec 15, 2025 | 0.5.4.post3 | 1.2.0rc3 | 0.11.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | release notes | - |
| v0.7.0.post1 | patch | - | 0.5.4.post3 | 1.2.0rc3 | 0.11.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | - | Post-train of v0.7.0: TensorRT-LLM pin advanced to 1.2.0rc3 (v0.7.0 shipped 1.2.0rc2). Same CUDA support as v0.7.0. |
| v0.7.0 | stable | Nov 26, 2025 | 0.5.4.post3 | 1.2.0rc2 | 0.11.0 | 0.8.0 / 0.8.0 / 0.8.0 | - | release notes | - |
| v0.6.1.post1 | patch | - | 0.5.3.post2 | 1.1.0rc5 | 0.11.0 | 0.6.0 / 0.6.0 / 0.6.0 | - | - | Post-train of v0.6.1: same backend pins as v0.6.1. Same CUDA support as v0.6.1. |
| v0.6.1 | patch | Nov 6, 2025 | 0.5.3.post2 | 1.1.0rc5 | 0.11.0 | 0.6.0 / 0.6.0 / 0.6.0 | - | release notes | - |
| v0.6.0 | stable | Oct 28, 2025 | 0.5.3.post2 | 1.1.0rc5 | 0.11.0 | 0.6.0 / 0.6.0 / 0.6.0 | - | release notes | Oldest release tracked on this page. |
Release highlights (stable releases):
- v1.4.0: Experimental cross-datacenter prefix routing and reservation replay in the Router, a vLLM-compatible generate token endpoint, tokenizer L1 prefix cache on by default, NIXL disaggregation for vLLM-Omni pipelines, and the Spica deployment simulator.
- v1.3.0: Tool-calling and reasoning overhaul, RL rollout serving, the largest Router buildout to date, SLA-driven Planner autoscaling, and production GPU Memory Service on Kubernetes.
- v1.2.0: DGD/DGDR v1beta1, CRTC as the default KV router, inter-pod GPU Memory Service, Dynamo Snapshot on CRI-O/OpenShift, and DeepSeek-V4 recipes on vLLM.
- v1.1.0: Resilient KV routing at scale, Anthropic Messages API support, performance modeling and offline replay, and the multimodal embedding cache.
- v1.0.0: First GA release: unified configuration, Kubernetes production readiness, multimodal serving, and the agents surface.
CUDA toolkit and minimum driver history
| Dynamo | Backend | CUDA Toolkit | Min Driver | Note |
|---|---|---|---|---|
| 1.4.2 | SGLang | 13.0 | 580.xx+ | - |
| 1.4.2 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.4.2 | vLLM | 13.0 | 580.xx+ | - |
| 1.4.1 | SGLang | 13.0 | 580.xx+ | - |
| 1.4.1 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.4.1 | vLLM | 13.0 | 580.xx+ | - |
| 1.4.0 | SGLang | 13.0 | 580.xx+ | - |
| 1.4.0 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.4.0 | vLLM | 13.0 | 580.xx+ | - |
| 1.3.1 | SGLang | 13.0 | 580.xx+ | - |
| 1.3.1 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.3.1 | vLLM | 13.0 | 580.xx+ | - |
| 1.3.0 | SGLang | 13.0 | 580.xx+ | - |
| 1.3.0 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.3.0 | vLLM | 13.0 | 580.xx+ | - |
| 1.2.1 | SGLang | 12.9 | 575.xx+ | - |
| 1.2.1 | SGLang | 13.0 | 580.xx+ | - |
| 1.2.1 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.2.1 | vLLM | 12.9 | 575.xx+ | - |
| 1.2.1 | vLLM | 13.0 | 580.xx+ | - |
| 1.2.0 | SGLang | 12.9 | 575.xx+ | - |
| 1.2.0 | SGLang | 13.0 | 580.xx+ | - |
| 1.2.0 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.2.0 | vLLM | 12.9 | 575.xx+ | - |
| 1.2.0 | vLLM | 13.0 | 580.xx+ | - |
| 1.1.1 | SGLang | 12.9 | 575.xx+ | - |
| 1.1.1 | SGLang | 13.0 | 580.xx+ | - |
| 1.1.1 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.1.1 | vLLM | 12.9 | 575.xx+ | - |
| 1.1.1 | vLLM | 13.0 | 580.xx+ | - |
| 1.1.0 | SGLang | 12.9 | 575.xx+ | - |
| 1.1.0 | SGLang | 13.0 | 580.xx+ | - |
| 1.1.0 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.1.0 | vLLM | 12.9 | 575.xx+ | - |
| 1.1.0 | vLLM | 13.0 | 580.xx+ | - |
| 1.0.2 | SGLang | 12.9 | 575.xx+ | - |
| 1.0.2 | SGLang | 13.0 | 580.xx+ | - |
| 1.0.2 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.0.2 | vLLM | 12.9 | 575.xx+ | - |
| 1.0.2 | vLLM | 13.0 | 580.xx+ | - |
| 1.0.1 | SGLang | 12.9 | 575.xx+ | - |
| 1.0.1 | SGLang | 13.0 | 580.xx+ | - |
| 1.0.1 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.0.1 | vLLM | 12.9 | 575.xx+ | - |
| 1.0.1 | vLLM | 13.0 | 580.xx+ | - |
| 1.0.0 | SGLang | 12.9 | 575.xx+ | - |
| 1.0.0 | SGLang | 13.0 | 580.xx+ | - |
| 1.0.0 | TensorRT-LLM | 13.1 | 580.xx+ | - |
| 1.0.0 | vLLM | 12.9 | 575.xx+ | - |
| 1.0.0 | vLLM | 13.0 | 580.xx+ | - |
| 0.9.1 | SGLang | 12.9 | 575.xx+ | - |
| 0.9.1 | TensorRT-LLM | 13.0 | 580.xx+ | - |
| 0.9.1 | vLLM | 12.9 | 575.xx+ | - |
| 0.9.0 | SGLang | 12.9 | 575.xx+ | - |
| 0.9.0 | TensorRT-LLM | 13.0 | 580.xx+ | - |
| 0.9.0 | vLLM | 12.9 | 575.xx+ | - |
| 0.8.1 | SGLang | 12.9 | 575.xx+ | - |
| 0.8.1 | SGLang | 13.0 | 580.xx+ | Experimental |
| 0.8.1 | TensorRT-LLM | 13.0 | 580.xx+ | - |
| 0.8.1 | vLLM | 12.9 | 575.xx+ | - |
| 0.8.1 | vLLM | 13.0 | 580.xx+ | Experimental |
| 0.8.0 | SGLang | 12.9 | 575.xx+ | - |
| 0.8.0 | SGLang | 13.0 | 580.xx+ | Experimental |
| 0.8.0 | TensorRT-LLM | 13.0 | 580.xx+ | - |
| 0.8.0 | vLLM | 12.9 | 575.xx+ | - |
| 0.8.0 | vLLM | 13.0 | 580.xx+ | Experimental |
| 0.7.1 | SGLang | 12.8 | 570.xx+ | - |
| 0.7.1 | TensorRT-LLM | 13.0 | 580.xx+ | - |
| 0.7.1 | vLLM | 12.9 | 575.xx+ | - |
| 0.7.0 | SGLang | 12.9 | 575.xx+ | - |
| 0.7.0 | TensorRT-LLM | 13.0 | 580.xx+ | - |
| 0.7.0 | vLLM | 12.8 | 570.xx+ | - |
- Patch versions (e.g. v0.8.1.post1, v0.7.0.post1) have the same CUDA support as their base version.
- Early access v1.1.0-dev.* images follow the same CUDA matrix as v1.0.2. The v1.2.0-deepseek-v4-dev.3 vLLM container is CUDA 13.0 multi-arch; the SGLang containers split by arch (CUDA 12.9 on amd64, CUDA 13.0 on arm64).
- Experimental CUDA 13 images are not published for all versions.
Feature support by backend (v1.4.2)
| Feature | SGLang | TensorRT-LLM | vLLM |
|---|---|---|---|
| Disaggregated Serving | Supported | Supported | Supported (Prefill/decode separation with NIXL KV transfer) |
| KV-Aware Routing | Supported | Supported | Supported |
| SLA-Based Planner | Supported | Supported | Supported |
| KV Block Manager | Experimental (Work in progress across all combinations) | Supported | Supported |
| Multimodal (Image) | Supported (KV-aware routing supported on Dynamo’s SGLang image for aggregated workers; a custom build without the hash-forwarding patch falls back to text-prefix routing. Separately, multimodal serving supports EPD, E/PD and E/P/D disaggregation (not traditional EP/D)) | Supported (Image URLs + pre-computed embeddings. Disagg: EP/D + E/P/D. KV-aware routing via dedicated MM Router Worker (requires KV event publishing)) | Supported (With KV-aware routing, image-aware routing on documented paths) |
| Multimodal (Video) | Supported | Not supported | Supported (Video input with frame sampling) |
| Multimodal (Audio) | Not supported | Not supported | Experimental (Qwen2-Audio, experimental) |
| Request Migration | Supported | Supported (Work in progress with multimodal) | Supported |
| Request Cancellation | Experimental (Remote-prefill-phase cancellation not supported in disaggregated mode) | Supported with caveat (Engine temporarily not notified of cancellations — resources for cancelled requests are not freed (known issue)) | Supported |
| LoRA | Experimental (Dynamic loading, discovery, and aggregated inference validated; unloading is implemented but not end-to-end tested; disaggregated serving not end-to-end validated) | Not supported | Supported (Dynamic load/unload; KV-aware routing supports adapter affinity) |
| Tool Calling | Supported | Supported | Supported |
| Speculative Decoding | Experimental (Code hooks exist; no examples or docs yet) | Supported | Supported (Eagle3) |
| GPU Memory Service | Supported (Weights and KV; upstream integration remains in progress) | Experimental (Weights only; multinode and upstream integration remain in progress) | Supported (Weights and KV; upstream integration remains in progress) |
| Shadow Engine Failover | Experimental (No KV-cache reuse or hardware fault tolerance) | Experimental (No KV-cache reuse or hardware fault tolerance) | Supported with caveat (Software-process failover only; no KV-cache reuse or hardware fault tolerance) |
| Dynamo Snapshot | Supported with caveat (Single-GPU supported; multi-GPU and multinode remain in progress) | Experimental (Single-GPU aggregated text-worker path only) | Supported with caveat (Single-GPU supported; multi-GPU is highly experimental and multinode remains in progress) |
Artifact inventory (v1.4.2)
| Category | Name | Description | Meta | Tags / install |
|---|---|---|---|---|
| container | vllm-runtime | vLLM backend runtime | vLLM v0.26.0 · CUDA 13.0 · AMD64/ARM64 | nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2; nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2-efa |
| container | sglang-runtime | SGLang backend runtime | SGLang v0.5.16 · CUDA 13.0 · AMD64/ARM64 | nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.4.2; nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.4.2-efa |
| container | tensorrtllm-runtime | TensorRT-LLM backend runtime | TRT-LLM v1.3.0rc22 · CUDA 13.1 · AMD64/ARM64 | nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.4.2; nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.4.2-efa |
| container | dynamo-frontend | OpenAI-compatible API gateway with Endpoint Prediction Protocol (EPP) | AMD64/ARM64 | nvcr.io/nvidia/ai-dynamo/dynamo-frontend:1.4.2 |
| container | dynamo-planner | Standalone Planner used by Profiler jobs and Planner pods | AMD64/ARM64 | nvcr.io/nvidia/ai-dynamo/dynamo-planner:1.4.2 |
| container | kubernetes-operator | Operator that manages Dynamo deployments and CRDs | AMD64/ARM64 | nvcr.io/nvidia/ai-dynamo/kubernetes-operator:1.4.2 |
| container | snapshot-agent (Preview) | Fast GPU worker recovery via CRIU | AMD64 | nvcr.io/nvidia/ai-dynamo/snapshot-agent:1.4.2 |
| wheel | ai-dynamo | Main package with backend integrations (vLLM, SGLang, TRT-LLM) | Python 3.10–3.12 · Linux (glibc v2.28+) | uv pip install ai-dynamo==1.4.2 |
| wheel | ai-dynamo-runtime | Core Python bindings for the Dynamo runtime | Python 3.10–3.12 · Linux (glibc v2.28+) | uv pip install ai-dynamo-runtime==1.4.2 |
| wheel | kvbm | KV Block Manager for disaggregated KV cache | Python 3.10–3.12 · Linux (glibc v2.28+) | uv pip install kvbm==1.4.2 |
| helm | dynamo-platform | Platform services (etcd, NATS) and the Dynamo Operator for a Dynamo cluster | - | helm install dynamo-platform https://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/dynamo-platform-1.4.2.tgz |
| helm | snapshot | Snapshot DaemonSet for fast GPU worker recovery | - | helm install snapshot https://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/snapshot-1.4.2.tgz |
| crate | dynamo-runtime | Core distributed runtime library | MSRV Rust v1.82 | cargo add dynamo-runtime@1.4.2 |
| crate | dynamo-llm | LLM inference engine | MSRV Rust v1.82 | cargo add dynamo-llm@1.4.2 |
| crate | dynamo-protocols | Async OpenAI-compatible API client | Independently versioned | cargo add dynamo-protocols@5.0.1 |
| crate | dynamo-async-openai (Deprecated) | Legacy OpenAI client; use dynamo-protocols | MSRV Rust v1.82 · final release | cargo add dynamo-async-openai@1.0.2 |
| crate | dynamo-parsers | Protocol parsers (SSE, JSON streaming) | Independently versioned | cargo add dynamo-parsers@7.0.1 |
| crate | dynamo-memory | Memory management utilities | MSRV Rust v1.82 | cargo add dynamo-memory@1.4.2 |
| crate | dynamo-config | Configuration management | MSRV Rust v1.82 | cargo add dynamo-config@1.2.1 |
| crate | dynamo-tokens | Tokenizer bindings for LLM inference | MSRV Rust v1.82 | cargo add dynamo-tokens@1.4.2 |
| crate | dynamo-tokenizers | Tokenizer library for LLM inference | Independently versioned | cargo add dynamo-tokenizers@1.5.4 |
| crate | dynamo-mocker | Inference engine simulator for benchmarking | MSRV Rust v1.82 | cargo add dynamo-mocker@1.4.2 |
| crate | dynamo-kv-router | KV-aware request routing library | MSRV Rust v1.82 | cargo add dynamo-kv-router@1.4.2 |
| crate | kvbm-logical | Logical layer for the KV Block Manager | MSRV Rust v1.82 | cargo add kvbm-logical@1.4.2 |
| crate | dynamo-kv-hashing | Request-to-lineage-hash contract for KV cache identity | MSRV Rust v1.82 | cargo add dynamo-kv-hashing@1.4.2 |
| crate | dynamo-data-gen | Schemas and primitives for Dynamo data generation and replay traces | MSRV Rust v1.82 | cargo add dynamo-data-gen@1.4.2 |
| crate | dynamo-rl | Dynamo RL worker discovery API | MSRV Rust v1.82 | cargo add dynamo-rl@1.4.2 |
| crate | dynamo-bench | Lightweight HTTP benchmarks for Dynamo endpoints | MSRV Rust v1.82 | cargo add dynamo-bench@1.4.2 |
| crate | dynamo-truthy | Canonical truthy/falsy boolean flag parsing | MSRV Rust v1.82 | cargo add dynamo-truthy@1.4.2 |
| crate | kvbm-common | Shared types for the KV Block Manager | MSRV Rust v1.82 | cargo add kvbm-common@1.4.2 |
| crate | kvbm-config | KVBM configuration for Tokio, Rayon, and Messenger runtimes | MSRV Rust v1.82 | cargo add kvbm-config@1.4.2 |
| crate | kvbm-kernels | CUDA kernels for the KV Block Manager | MSRV Rust v1.82 | cargo add kvbm-kernels@1.4.2 |
| crate | kvbm-physical | Physical block layer for the KV Block Manager | MSRV Rust v1.82 | cargo add kvbm-physical@1.4.2 |
| crate | kvbm-engine | Distributed coordination primitives for KVBM | MSRV Rust v1.82 | cargo add kvbm-engine@1.4.2 |
| crate | dynamo-renderer | Chat-template rendering used by the Dynamo Frontend | Independently versioned | cargo add dynamo-renderer@4.0.0 |
| crate | dynamo-parsers-v2 | Successor parser line to dynamo-parsers, consumed by the Frontend | Independently versioned | cargo add dynamo-parsers-v2@0.1.23 |
| crate | fastokens | Rust BPE tokenizer backend consumed by the Frontend | Independently versioned | cargo add fastokens@0.2.0 |
Known artifact issues
| Release | Artifact | Issue | Status |
|---|---|---|---|
| v0.9.0 | dynamo-platform-0.9.0 | Helm chart sets operator image to 0.7.1 instead of 0.9.0. | Fixed in v0.9.0.post1 |
| v0.8.1 | vllm-runtime:0.8.1-cuda13 | Container fails to launch. | Known issue |
| v0.8.1 | sglang-runtime:0.8.1-cuda13, vllm-runtime:0.8.1-cuda13 | Multimodality not expected to work on ARM64. Works on AMD64. | Known limitation |
| v0.8.0 | sglang-runtime:0.8.0-cuda13 | CuDNN installation issue caused PyTorch v2.9.1 compatibility problems with nn.Conv3d — performance degradation and excessive memory usage in multimodal workloads. | Fixed in v0.8.1 (#5461) |
Crates: first published version on crates.io
| Crate | First version | Date |
|---|---|---|
| dynamo-runtime | 0.1.0 | 2025-03-18 |
| dynamo-llm | 0.2.0 | 2025-05-01 |
| dynamo-async-openai | 0.4.1 | 2025-08-27 |
| dynamo-parsers | 0.5.0 | 2025-09-18 |
| dynamo-memory | 0.8.0 | 2026-01-15 |
| dynamo-config | 0.8.0 | 2026-01-15 |
| dynamo-tokens | 0.9.0 | 2026-02-12 |
| dynamo-mocker | 1.0.0 | 2026-03-13 |
| dynamo-kv-router | 1.0.0 | 2026-03-13 |
| dynamo-protocols | 1.1.0 | 2026-05-04 |
| dynamo-tokenizers | 1.2.0 | 2026-06-02 |
Model early-access builds
| Model | Tag | Release line | Runtimes | Shipped | GA path | Status | Coverage (images / wheels / helm / crates) |
|---|---|---|---|---|---|---|---|
| Inkling | 1.4.0-inkling-dev.1 | v1.4.0 | sglang-runtime | Jul 17, 2026 | Dev-only · v1.4.0 line | First build on the v1.4.0 line; targets the next stable release. | yes / no / no / no |
| GLM-5.2 | 1.3.0-glm-5.2-dev.1 | v1.3.0 | sglang-runtime | Jul 20, 2026 | Dev-only | Container carries SGLang cherry-picks (stability, config parsing, model support) opened upstream but not yet in a released SGLang. | yes / no / no / no |
| MiniMax-M3 | 1.3.0-minimax-m3-dev.1 | v1.3.0 | vllm-runtime, sglang-runtime, tensorrtllm-runtime | Jun 12, 2026 | Promoted → :1.3.0 | Dynamo changes and the M2 tool-calling fix are in release/1.3.0; the recipes run on the stock :1.3.0 containers. | yes / no / no / no |
| DeepSeek-V4 | 1.3.0-deepseek-v4-dev.1 | v1.3.0 | tensorrtllm-runtime | Jun 6, 2026 | Recipe in v1.3.0 | DeepSeek-V4 Flash and Pro recipes ship in v1.3.0 on the standard TensorRT-LLM release container. | yes / no / no / no |
| Nemotron-3-Ultra | 1.3.0-nemotron-ultra-dev.1 | v1.3.0 | vllm-runtime | Jun 5, 2026 | Dev-only | Four un-upstreamed vLLM patches; requires pinned flags VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernel and —no-enable-flashinfer-autotune. | yes / no / no / no |
| Nemotron-3-Super | 1.3.0-nemotron-super-dev.1 | v1.3.0 | vllm-runtime | Jun 4, 2026 | Dev-only | Requires the dedicated vllm-runtime:1.3.0-nemotron-super-dev.1 image; the model-specific vLLM patches are not in the v1.3.0 release container. | yes / no / no / no |
| Kimi-K2.6 | 1.3.0-kimi-k2.6-dev.1 | v1.3.0 | vllm-runtime | Jun 4, 2026 | Promoted → :1.3.0 | The build’s only container patch is in vLLM v0.23.0; the recipes run on the stock vllm-runtime:1.3.0. | yes / no / no / no |
| Cosmos-3 | 1.3.0-cosmos3-dev.1 | v1.3.0 | vllm-runtime | Jun 1, 2026 | Dev-only | Dynamo #10132 (Cosmos3 support in the vLLM-Omni backend) is open, not merged — v1.3.0 containers cannot run Cosmos3. | yes / no / no / no |
| DeepSeek-V4 preview | 1.2.0-deepseek-v4-dev.3 | v1.2.0 | vllm-runtime, sglang-runtime | May 9, 2026 | Superseded — recipe in v1.3.0 | Blackwell (B200 + GB200) preview; per-arch/CUDA tags (e.g. vllm-runtime:1.2.0-deepseek-v4-cuda13-dev.3). Superseded by the v1.3.0 recipe. | yes / no / no / no |
| DeepSeek-V4 preview | 1.2.0-deepseek-v4-dev.2 | v1.2.0 | vllm-runtime, sglang-runtime | May 1, 2026 | Superseded — recipe in v1.3.0 | Blackwell preview on vLLM v0.20.0 (native DSv4 support); superseded by dev.3. | yes / no / no / no |
| DeepSeek-V4 preview | 1.2.0-sglang-deepseek-v4-dev.1 | v1.2.0 | sglang-runtime | Apr 25, 2026 | Superseded — recipe in v1.3.0 | Earliest DSv4 preview (SGLang, B200 only); superseded by dev.2/dev.3. | yes / no / no / no |
Platform-preview artifact coverage
| Preview | Images | Wheels | Helm | Crates |
|---|---|---|---|---|
| v1.3.0-dev.1 | yes | yes | yes | yes |
| v1.1.0-dev.3 | yes | yes | no | no |
| v1.1.0-dev.2 | yes | yes | no | no |
| v1.1.0-dev.1 | yes | yes | yes | no |
Platform support
- GPU architectures: Blackwell, Hopper, Ada Lovelace, Ampere
- OS: Ubuntu 24.04 (x86_64, ARM64) — Containers and wheels
- OS: Ubuntu 22.04 (x86_64) — Wheels only
- CSP: AWS — Amazon Linux 2023 (x86_64) — Containers and wheels
- CPU architectures: x86_64, ARM64 (Ubuntu 24.04 only)
- Wheels: Wheels are built in a manylinux_2_28 environment (AlmaLinux 8, glibc 2.28+) and validated on Ubuntu 22.04 and 24.04. They install on any Linux distribution with glibc 2.28+ (Debian 11+, RHEL 9, etc.), but only Ubuntu 22.04/24.04 are officially verified.
Release statistics
| Release | PRs | Contributors | First-time contributors | Breaking changes | Known issues |
|---|---|---|---|---|---|
| v1.4.0 | 640 | 127 | 29 | 51 | 19 |
| v1.3.0 | 930 | 125 | 24 | 24 | 10 |
| v1.2.0 | 603 | 82 | - | 5 | 11 |
| v1.1.0 | 896 | 113 | - | 8 | 20 |
| v1.0.0 | - | 90 | 34 | 41 | 14 |
| v0.9.0 | - | - | 14 | 1 | 13 |
| v0.8.0 | - | - | 20 | 0 | 14 |
| v0.7.0 | - | - | 2 | 0 | 7 |
| v0.6.0 | - | - | 4 | 0 | 3 |
Nightlies
ai-dynamo and ai-dynamo-runtime nightly builds from main publish wheels tagged *.devYYYYMMDD (since Apr 24, 2026); kvbm joined the nightly train on Aug 2, 2026. Install with pip or uv using --pre and the NVIDIA extra-index pattern shown above. Runtime containers publish to the *-runtime-nightly repositories on NGC, under a dated YYYYMMDD-<shortsha> tag plus a rolling latest tag.
| Version | Date | Packages | Notes |
|---|---|---|---|
| 1.5.0.dev20260831 | Aug 31, 2026 | ai-dynamo, ai-dynamo-runtime, kvbm | - |
| 1.5.0.dev20260830 | Aug 30, 2026 | ai-dynamo, ai-dynamo-runtime, kvbm | - |
| 1.5.0.dev20260829 | Aug 29, 2026 | ai-dynamo, ai-dynamo-runtime, kvbm | - |