Releases (machine-readable)

Generated release, compatibility, and artifact data for agents and automation
View as Markdown

This page is a plain-markdown rendering of components/releases.data.ts, the single source of truth behind Compatibility, Release Artifacts, Model Early Access Builds, and the Release Notes timeline. It is regenerated by scripts/gen_llms_tables.py at every release bump — do not edit the tables below by hand. Append .md to this page’s URL for a clean markdown export. Fetch the same data as JSON or an Atom feed from the docs-website branch, which the publish pipeline regenerates on every docs publish; the files are no longer committed to main.

Current stable release: v1.4.2 (Aug 28, 2026; container tag 1.4.2, wheel version 1.4.2).

Releases

VersionKindDateSGLangTensorRT-LLMvLLMNIXL (SGL / TRT / vLLM)UCXNotesDelta
main (ToT)development head-0.5.181.3.0rc250.28.01.4.0 / 1.3.1 / 1.3.2---
v1.4.2patchAug 28, 20260.5.161.3.0rc220.26.01.3.0 / 1.3.1 / 1.3.21.21.xrelease notesPatch release and the first Dynamo Enterprise release: a curated set of release artifacts publishes under the -enterprise suffix on NGC, eligible for NVIDIA Enterprise Support, with no functional or binary differences from the open-source artifacts. Fixes NIXL loader-path resolution in the Frontend and SGLang Runtime images, removes the unused Nsight EFA metrics plugin, and tightens dependency pins (pillow v12.3.0 floor, plotext below v6, EFA Installer v1.50). Backend pins are unchanged from v1.4.0.
v1.4.1patchAug 21, 20260.5.161.3.0rc220.26.01.3.0 / 1.3.1 / 1.3.21.21.xrelease notesPatch release. Adds the classify and pooling endpoints, forwards logprob_token_ids through the OpenAI frontend, reconciles request-path overload marks in the Router, and fixes NIXL writable buffers for vLLM. All three Go modules move to Go 1.26.6 with aligned x/net and grpc. Backend pins are unchanged from v1.4.0.
v1.4.0stableAug 14, 20260.5.161.3.0rc220.26.01.3.0 / 1.3.1 / 1.3.21.21.xrelease notesAudit subsystem migrated into request trace (DYN_AUDIT_* honored as legacy aliases); HTTP header capture in trace records is an explicit fail-closed allowlist; deprecated multimodal worker flags and vLLM worker-role flags removed; runtime images no longer bundle software video decoders (H.264/H.265 decodes via NVDEC); UCX 1.21.x.
v1.3.1patchAug 5, 20260.5.141.3.0rc190.23.01.3.2 / 1.0.1 / 1.1.01.20.xrelease notesPatch release. Fixes disaggregated SGLang serving over AWS EFA on GB200: the SGLang EFA runtime moves to NIXL 1.3.2 and all three EFA images to EFA Installer 1.49.0. Backend pins are otherwise unchanged from v1.3.0.
v1.3.0stableJul 20, 20260.5.141.3.0rc190.23.01.3.0 / 1.0.1 / 1.1.01.20.xrelease notesCUDA 12 container images discontinued; EFA variants retagged from -efa-amd64 to -efa (the images were already multi-arch — the old suffix was misleading); GA wheels published as 1.3.0.post1 (containers stay :1.3.0); UCX 1.20.x.
v1.3.0-dev.1platform-previewJun 9, 20260.5.12.post11.3.0rc170.22.01.0.1 / 0.10.1 / 1.1.0-release notesFull-platform preview of v1.3.0: complete runtime matrix, wheels on pypi.nvidia.com, crates, and Helm charts. Superseded by v1.3.0 GA.
v1.2.1patchJun 13, 20260.5.111.3.0rc140.20.11.0.1 / 0.10.1 / 0.10.1-release notesPatch release. Same backend pins as v1.2.0.
v1.2.0stableJun 2, 20260.5.111.3.0rc140.20.11.0.1 / 0.10.1 / 0.10.11.20.0release notes603 PRs from 82 authors. DGD/DGDR promoted to v1beta1; CRTC default approximate KV router; inter-pod GMS sidecar; Dynamo Snapshot on CRI-O / OpenShift; UCX 1.20.0.
v1.2.0-deepseek-v4-dev.3model-buildMay 9, 2026upstream DSv4 preview-0.20.1- / - / 0.10.1-release notesDeepSeek-V4 Blackwell preview; vLLM + SGLang containers only.
v1.2.0-deepseek-v4-dev.2model-buildMay 1, 2026upstream DSv4 preview-0.20.0- / - / 0.10.1-release notesDeepSeek-V4 Blackwell preview; vLLM + SGLang containers only.
v1.1.1patchMay 5, 20260.5.10.post11.3.0rc110.19.01.0.1 / 0.10.1 / 0.10.1-release notesPatch release. Same backend pins as v1.1.0.
v1.1.0stableMay 1, 20260.5.10.post11.3.0rc110.19.01.0.1 / 0.10.1 / 0.10.11.20release notesPlanner split into its own dynamo-planner image (artifact boundary change). First 1.y.z publication of dynamo-protocols on crates.io; dynamo-async-openai deprecated at final 1.0.2.
v1.1.0-dev.3platform-previewApr 18, 20260.5.10.post11.3.0rc110.19.01.0.1 / 0.10.1 / 0.10.1-release notesPartial platform preview: TRT-LLM runtime image + wheels only.
v1.1.0-dev.2platform-previewApr 9, 20260.5.91.3.0rc90.19.01.0.1 / 0.10.1 / 0.10.1-release notesPartial platform preview: SGLang + TRT-LLM runtime images + wheels.
v1.1.0-dev.1platform-previewMar 17, 20260.5.91.3.0rc5.post10.17.11.0.1 / 0.10.1 / 0.10.1-release notesPlatform preview: runtime matrix, wheels on pypi.nvidia.com, Helm charts.
v1.0.2patchApr 22, 20260.5.91.3.0rc5.post10.16.00.10.1 / 0.10.1 / 0.10.1-release notesNo artifact additions or removals versus v1.0.0.
v1.0.1patchMar 16, 20260.5.91.3.0rc5.post10.16.00.10.1 / 0.10.1 / 0.10.1-release notesNo artifact additions or removals versus v1.0.0.
v1.0.0stableMar 12, 20260.5.91.3.0rc5.post10.16.00.10.1 / 0.10.1 / 0.10.1-release notessnapshot-agent image and EFA variants for vLLM and TensorRT-LLM. First publish of dynamo-mocker and dynamo-kv-router crates. snapshot Helm chart added (preview); deprecated dynamo-crds dropped from the publish stream.
v0.9.1patchMar 4, 20260.5.81.3.0rc30.14.10.9.0 / 0.9.0 / 0.9.0-release notesNo artifact additions or removals versus v0.9.0.
v0.9.0stableFeb 11, 20260.5.81.3.0rc10.14.10.9.0 / 0.9.0 / 0.9.0-release notesFirst publish of dynamo-tokens crate. Deprecated dynamo-graph Helm chart dropped from the publish stream.
v0.8.1.post3patch-0.5.6.post21.2.0rc6.post30.12.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post3. Same CUDA support as v0.8.1.
v0.8.1.post2patch-0.5.6.post21.2.0rc6.post20.12.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post2. Same CUDA support as v0.8.1.
v0.8.1.post1patch-0.5.6.post21.2.0rc6.post10.12.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.8.1: republished the TensorRT-LLM runtime image and PyPI wheels only, with TRT-LLM pinned to 1.2.0rc6.post1. Same CUDA support as v0.8.1.
v0.8.1patchJan 23, 20260.5.6.post21.2.0rc6.post10.12.00.8.0 / 0.8.0 / 0.8.0-release notesPost trains .post1/.post2/.post3 republished the TRT-LLM runtime image and PyPI wheels only; each carried a distinct TRT-LLM pin (see the v0.8.1.post1/.post2/.post3 rows).
v0.8.0stableJan 15, 20260.5.6.post21.2.0rc6.post10.12.00.8.0 / 0.8.0 / 0.8.0-release notesdynamo-frontend image and CUDA 13 variants for vLLM and SGLang. First publish of dynamo-memory and dynamo-config crates.
v0.7.1patchDec 15, 20250.5.4.post31.2.0rc30.11.00.8.0 / 0.8.0 / 0.8.0-release notes-
v0.7.0.post1patch-0.5.4.post31.2.0rc30.11.00.8.0 / 0.8.0 / 0.8.0--Post-train of v0.7.0: TensorRT-LLM pin advanced to 1.2.0rc3 (v0.7.0 shipped 1.2.0rc2). Same CUDA support as v0.7.0.
v0.7.0stableNov 26, 20250.5.4.post31.2.0rc20.11.00.8.0 / 0.8.0 / 0.8.0-release notes-
v0.6.1.post1patch-0.5.3.post21.1.0rc50.11.00.6.0 / 0.6.0 / 0.6.0--Post-train of v0.6.1: same backend pins as v0.6.1. Same CUDA support as v0.6.1.
v0.6.1patchNov 6, 20250.5.3.post21.1.0rc50.11.00.6.0 / 0.6.0 / 0.6.0-release notes-
v0.6.0stableOct 28, 20250.5.3.post21.1.0rc50.11.00.6.0 / 0.6.0 / 0.6.0-release notesOldest release tracked on this page.

Release highlights (stable releases):

  • v1.4.0: Experimental cross-datacenter prefix routing and reservation replay in the Router, a vLLM-compatible generate token endpoint, tokenizer L1 prefix cache on by default, NIXL disaggregation for vLLM-Omni pipelines, and the Spica deployment simulator.
  • v1.3.0: Tool-calling and reasoning overhaul, RL rollout serving, the largest Router buildout to date, SLA-driven Planner autoscaling, and production GPU Memory Service on Kubernetes.
  • v1.2.0: DGD/DGDR v1beta1, CRTC as the default KV router, inter-pod GPU Memory Service, Dynamo Snapshot on CRI-O/OpenShift, and DeepSeek-V4 recipes on vLLM.
  • v1.1.0: Resilient KV routing at scale, Anthropic Messages API support, performance modeling and offline replay, and the multimodal embedding cache.
  • v1.0.0: First GA release: unified configuration, Kubernetes production readiness, multimodal serving, and the agents surface.

CUDA toolkit and minimum driver history

DynamoBackendCUDA ToolkitMin DriverNote
1.4.2SGLang13.0580.xx+-
1.4.2TensorRT-LLM13.1580.xx+-
1.4.2vLLM13.0580.xx+-
1.4.1SGLang13.0580.xx+-
1.4.1TensorRT-LLM13.1580.xx+-
1.4.1vLLM13.0580.xx+-
1.4.0SGLang13.0580.xx+-
1.4.0TensorRT-LLM13.1580.xx+-
1.4.0vLLM13.0580.xx+-
1.3.1SGLang13.0580.xx+-
1.3.1TensorRT-LLM13.1580.xx+-
1.3.1vLLM13.0580.xx+-
1.3.0SGLang13.0580.xx+-
1.3.0TensorRT-LLM13.1580.xx+-
1.3.0vLLM13.0580.xx+-
1.2.1SGLang12.9575.xx+-
1.2.1SGLang13.0580.xx+-
1.2.1TensorRT-LLM13.1580.xx+-
1.2.1vLLM12.9575.xx+-
1.2.1vLLM13.0580.xx+-
1.2.0SGLang12.9575.xx+-
1.2.0SGLang13.0580.xx+-
1.2.0TensorRT-LLM13.1580.xx+-
1.2.0vLLM12.9575.xx+-
1.2.0vLLM13.0580.xx+-
1.1.1SGLang12.9575.xx+-
1.1.1SGLang13.0580.xx+-
1.1.1TensorRT-LLM13.1580.xx+-
1.1.1vLLM12.9575.xx+-
1.1.1vLLM13.0580.xx+-
1.1.0SGLang12.9575.xx+-
1.1.0SGLang13.0580.xx+-
1.1.0TensorRT-LLM13.1580.xx+-
1.1.0vLLM12.9575.xx+-
1.1.0vLLM13.0580.xx+-
1.0.2SGLang12.9575.xx+-
1.0.2SGLang13.0580.xx+-
1.0.2TensorRT-LLM13.1580.xx+-
1.0.2vLLM12.9575.xx+-
1.0.2vLLM13.0580.xx+-
1.0.1SGLang12.9575.xx+-
1.0.1SGLang13.0580.xx+-
1.0.1TensorRT-LLM13.1580.xx+-
1.0.1vLLM12.9575.xx+-
1.0.1vLLM13.0580.xx+-
1.0.0SGLang12.9575.xx+-
1.0.0SGLang13.0580.xx+-
1.0.0TensorRT-LLM13.1580.xx+-
1.0.0vLLM12.9575.xx+-
1.0.0vLLM13.0580.xx+-
0.9.1SGLang12.9575.xx+-
0.9.1TensorRT-LLM13.0580.xx+-
0.9.1vLLM12.9575.xx+-
0.9.0SGLang12.9575.xx+-
0.9.0TensorRT-LLM13.0580.xx+-
0.9.0vLLM12.9575.xx+-
0.8.1SGLang12.9575.xx+-
0.8.1SGLang13.0580.xx+Experimental
0.8.1TensorRT-LLM13.0580.xx+-
0.8.1vLLM12.9575.xx+-
0.8.1vLLM13.0580.xx+Experimental
0.8.0SGLang12.9575.xx+-
0.8.0SGLang13.0580.xx+Experimental
0.8.0TensorRT-LLM13.0580.xx+-
0.8.0vLLM12.9575.xx+-
0.8.0vLLM13.0580.xx+Experimental
0.7.1SGLang12.8570.xx+-
0.7.1TensorRT-LLM13.0580.xx+-
0.7.1vLLM12.9575.xx+-
0.7.0SGLang12.9575.xx+-
0.7.0TensorRT-LLM13.0580.xx+-
0.7.0vLLM12.8570.xx+-
  • Patch versions (e.g. v0.8.1.post1, v0.7.0.post1) have the same CUDA support as their base version.
  • Early access v1.1.0-dev.* images follow the same CUDA matrix as v1.0.2. The v1.2.0-deepseek-v4-dev.3 vLLM container is CUDA 13.0 multi-arch; the SGLang containers split by arch (CUDA 12.9 on amd64, CUDA 13.0 on arm64).
  • Experimental CUDA 13 images are not published for all versions.

Feature support by backend (v1.4.2)

FeatureSGLangTensorRT-LLMvLLM
Disaggregated ServingSupportedSupportedSupported (Prefill/decode separation with NIXL KV transfer)
KV-Aware RoutingSupportedSupportedSupported
SLA-Based PlannerSupportedSupportedSupported
KV Block ManagerExperimental (Work in progress across all combinations)SupportedSupported
Multimodal (Image)Supported (KV-aware routing supported on Dynamo’s SGLang image for aggregated workers; a custom build without the hash-forwarding patch falls back to text-prefix routing. Separately, multimodal serving supports EPD, E/PD and E/P/D disaggregation (not traditional EP/D))Supported (Image URLs + pre-computed embeddings. Disagg: EP/D + E/P/D. KV-aware routing via dedicated MM Router Worker (requires KV event publishing))Supported (With KV-aware routing, image-aware routing on documented paths)
Multimodal (Video)SupportedNot supportedSupported (Video input with frame sampling)
Multimodal (Audio)Not supportedNot supportedExperimental (Qwen2-Audio, experimental)
Request MigrationSupportedSupported (Work in progress with multimodal)Supported
Request CancellationExperimental (Remote-prefill-phase cancellation not supported in disaggregated mode)Supported with caveat (Engine temporarily not notified of cancellations — resources for cancelled requests are not freed (known issue))Supported
LoRAExperimental (Dynamic loading, discovery, and aggregated inference validated; unloading is implemented but not end-to-end tested; disaggregated serving not end-to-end validated)Not supportedSupported (Dynamic load/unload; KV-aware routing supports adapter affinity)
Tool CallingSupportedSupportedSupported
Speculative DecodingExperimental (Code hooks exist; no examples or docs yet)SupportedSupported (Eagle3)
GPU Memory ServiceSupported (Weights and KV; upstream integration remains in progress)Experimental (Weights only; multinode and upstream integration remain in progress)Supported (Weights and KV; upstream integration remains in progress)
Shadow Engine FailoverExperimental (No KV-cache reuse or hardware fault tolerance)Experimental (No KV-cache reuse or hardware fault tolerance)Supported with caveat (Software-process failover only; no KV-cache reuse or hardware fault tolerance)
Dynamo SnapshotSupported with caveat (Single-GPU supported; multi-GPU and multinode remain in progress)Experimental (Single-GPU aggregated text-worker path only)Supported with caveat (Single-GPU supported; multi-GPU is highly experimental and multinode remains in progress)

Artifact inventory (v1.4.2)

CategoryNameDescriptionMetaTags / install
containervllm-runtimevLLM backend runtimevLLM v0.26.0 · CUDA 13.0 · AMD64/ARM64nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2; nvcr.io/nvidia/ai-dynamo/vllm-runtime:1.4.2-efa
containersglang-runtimeSGLang backend runtimeSGLang v0.5.16 · CUDA 13.0 · AMD64/ARM64nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.4.2; nvcr.io/nvidia/ai-dynamo/sglang-runtime:1.4.2-efa
containertensorrtllm-runtimeTensorRT-LLM backend runtimeTRT-LLM v1.3.0rc22 · CUDA 13.1 · AMD64/ARM64nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.4.2; nvcr.io/nvidia/ai-dynamo/tensorrtllm-runtime:1.4.2-efa
containerdynamo-frontendOpenAI-compatible API gateway with Endpoint Prediction Protocol (EPP)AMD64/ARM64nvcr.io/nvidia/ai-dynamo/dynamo-frontend:1.4.2
containerdynamo-plannerStandalone Planner used by Profiler jobs and Planner podsAMD64/ARM64nvcr.io/nvidia/ai-dynamo/dynamo-planner:1.4.2
containerkubernetes-operatorOperator that manages Dynamo deployments and CRDsAMD64/ARM64nvcr.io/nvidia/ai-dynamo/kubernetes-operator:1.4.2
containersnapshot-agent (Preview)Fast GPU worker recovery via CRIUAMD64nvcr.io/nvidia/ai-dynamo/snapshot-agent:1.4.2
wheelai-dynamoMain package with backend integrations (vLLM, SGLang, TRT-LLM)Python 3.10–3.12 · Linux (glibc v2.28+)uv pip install ai-dynamo==1.4.2
wheelai-dynamo-runtimeCore Python bindings for the Dynamo runtimePython 3.10–3.12 · Linux (glibc v2.28+)uv pip install ai-dynamo-runtime==1.4.2
wheelkvbmKV Block Manager for disaggregated KV cachePython 3.10–3.12 · Linux (glibc v2.28+)uv pip install kvbm==1.4.2
helmdynamo-platformPlatform services (etcd, NATS) and the Dynamo Operator for a Dynamo cluster-helm install dynamo-platform https://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/dynamo-platform-1.4.2.tgz
helmsnapshotSnapshot DaemonSet for fast GPU worker recovery-helm install snapshot https://helm.ngc.nvidia.com/nvidia/ai-dynamo/charts/snapshot-1.4.2.tgz
cratedynamo-runtimeCore distributed runtime libraryMSRV Rust v1.82cargo add dynamo-runtime@1.4.2
cratedynamo-llmLLM inference engineMSRV Rust v1.82cargo add dynamo-llm@1.4.2
cratedynamo-protocolsAsync OpenAI-compatible API clientIndependently versionedcargo add dynamo-protocols@5.0.1
cratedynamo-async-openai (Deprecated)Legacy OpenAI client; use dynamo-protocolsMSRV Rust v1.82 · final releasecargo add dynamo-async-openai@1.0.2
cratedynamo-parsersProtocol parsers (SSE, JSON streaming)Independently versionedcargo add dynamo-parsers@7.0.1
cratedynamo-memoryMemory management utilitiesMSRV Rust v1.82cargo add dynamo-memory@1.4.2
cratedynamo-configConfiguration managementMSRV Rust v1.82cargo add dynamo-config@1.2.1
cratedynamo-tokensTokenizer bindings for LLM inferenceMSRV Rust v1.82cargo add dynamo-tokens@1.4.2
cratedynamo-tokenizersTokenizer library for LLM inferenceIndependently versionedcargo add dynamo-tokenizers@1.5.4
cratedynamo-mockerInference engine simulator for benchmarkingMSRV Rust v1.82cargo add dynamo-mocker@1.4.2
cratedynamo-kv-routerKV-aware request routing libraryMSRV Rust v1.82cargo add dynamo-kv-router@1.4.2
cratekvbm-logicalLogical layer for the KV Block ManagerMSRV Rust v1.82cargo add kvbm-logical@1.4.2
cratedynamo-kv-hashingRequest-to-lineage-hash contract for KV cache identityMSRV Rust v1.82cargo add dynamo-kv-hashing@1.4.2
cratedynamo-data-genSchemas and primitives for Dynamo data generation and replay tracesMSRV Rust v1.82cargo add dynamo-data-gen@1.4.2
cratedynamo-rlDynamo RL worker discovery APIMSRV Rust v1.82cargo add dynamo-rl@1.4.2
cratedynamo-benchLightweight HTTP benchmarks for Dynamo endpointsMSRV Rust v1.82cargo add dynamo-bench@1.4.2
cratedynamo-truthyCanonical truthy/falsy boolean flag parsingMSRV Rust v1.82cargo add dynamo-truthy@1.4.2
cratekvbm-commonShared types for the KV Block ManagerMSRV Rust v1.82cargo add kvbm-common@1.4.2
cratekvbm-configKVBM configuration for Tokio, Rayon, and Messenger runtimesMSRV Rust v1.82cargo add kvbm-config@1.4.2
cratekvbm-kernelsCUDA kernels for the KV Block ManagerMSRV Rust v1.82cargo add kvbm-kernels@1.4.2
cratekvbm-physicalPhysical block layer for the KV Block ManagerMSRV Rust v1.82cargo add kvbm-physical@1.4.2
cratekvbm-engineDistributed coordination primitives for KVBMMSRV Rust v1.82cargo add kvbm-engine@1.4.2
cratedynamo-rendererChat-template rendering used by the Dynamo FrontendIndependently versionedcargo add dynamo-renderer@4.0.0
cratedynamo-parsers-v2Successor parser line to dynamo-parsers, consumed by the FrontendIndependently versionedcargo add dynamo-parsers-v2@0.1.23
cratefastokensRust BPE tokenizer backend consumed by the FrontendIndependently versionedcargo add fastokens@0.2.0

Known artifact issues

ReleaseArtifactIssueStatus
v0.9.0dynamo-platform-0.9.0Helm chart sets operator image to 0.7.1 instead of 0.9.0.Fixed in v0.9.0.post1
v0.8.1vllm-runtime:0.8.1-cuda13Container fails to launch.Known issue
v0.8.1sglang-runtime:0.8.1-cuda13, vllm-runtime:0.8.1-cuda13Multimodality not expected to work on ARM64. Works on AMD64.Known limitation
v0.8.0sglang-runtime:0.8.0-cuda13CuDNN installation issue caused PyTorch v2.9.1 compatibility problems with nn.Conv3d — performance degradation and excessive memory usage in multimodal workloads.Fixed in v0.8.1 (#5461)

Crates: first published version on crates.io

CrateFirst versionDate
dynamo-runtime0.1.02025-03-18
dynamo-llm0.2.02025-05-01
dynamo-async-openai0.4.12025-08-27
dynamo-parsers0.5.02025-09-18
dynamo-memory0.8.02026-01-15
dynamo-config0.8.02026-01-15
dynamo-tokens0.9.02026-02-12
dynamo-mocker1.0.02026-03-13
dynamo-kv-router1.0.02026-03-13
dynamo-protocols1.1.02026-05-04
dynamo-tokenizers1.2.02026-06-02

Model early-access builds

ModelTagRelease lineRuntimesShippedGA pathStatusCoverage (images / wheels / helm / crates)
Inkling1.4.0-inkling-dev.1v1.4.0sglang-runtimeJul 17, 2026Dev-only · v1.4.0 lineFirst build on the v1.4.0 line; targets the next stable release.yes / no / no / no
GLM-5.21.3.0-glm-5.2-dev.1v1.3.0sglang-runtimeJul 20, 2026Dev-onlyContainer carries SGLang cherry-picks (stability, config parsing, model support) opened upstream but not yet in a released SGLang.yes / no / no / no
MiniMax-M31.3.0-minimax-m3-dev.1v1.3.0vllm-runtime, sglang-runtime, tensorrtllm-runtimeJun 12, 2026Promoted → :1.3.0Dynamo changes and the M2 tool-calling fix are in release/1.3.0; the recipes run on the stock :1.3.0 containers.yes / no / no / no
DeepSeek-V41.3.0-deepseek-v4-dev.1v1.3.0tensorrtllm-runtimeJun 6, 2026Recipe in v1.3.0DeepSeek-V4 Flash and Pro recipes ship in v1.3.0 on the standard TensorRT-LLM release container.yes / no / no / no
Nemotron-3-Ultra1.3.0-nemotron-ultra-dev.1v1.3.0vllm-runtimeJun 5, 2026Dev-onlyFour un-upstreamed vLLM patches; requires pinned flags VLLM_DISABLED_KERNELS=FlashInferFP8ScaledMMLinearKernel and —no-enable-flashinfer-autotune.yes / no / no / no
Nemotron-3-Super1.3.0-nemotron-super-dev.1v1.3.0vllm-runtimeJun 4, 2026Dev-onlyRequires the dedicated vllm-runtime:1.3.0-nemotron-super-dev.1 image; the model-specific vLLM patches are not in the v1.3.0 release container.yes / no / no / no
Kimi-K2.61.3.0-kimi-k2.6-dev.1v1.3.0vllm-runtimeJun 4, 2026Promoted → :1.3.0The build’s only container patch is in vLLM v0.23.0; the recipes run on the stock vllm-runtime:1.3.0.yes / no / no / no
Cosmos-31.3.0-cosmos3-dev.1v1.3.0vllm-runtimeJun 1, 2026Dev-onlyDynamo #10132 (Cosmos3 support in the vLLM-Omni backend) is open, not merged — v1.3.0 containers cannot run Cosmos3.yes / no / no / no
DeepSeek-V4 preview1.2.0-deepseek-v4-dev.3v1.2.0vllm-runtime, sglang-runtimeMay 9, 2026Superseded — recipe in v1.3.0Blackwell (B200 + GB200) preview; per-arch/CUDA tags (e.g. vllm-runtime:1.2.0-deepseek-v4-cuda13-dev.3). Superseded by the v1.3.0 recipe.yes / no / no / no
DeepSeek-V4 preview1.2.0-deepseek-v4-dev.2v1.2.0vllm-runtime, sglang-runtimeMay 1, 2026Superseded — recipe in v1.3.0Blackwell preview on vLLM v0.20.0 (native DSv4 support); superseded by dev.3.yes / no / no / no
DeepSeek-V4 preview1.2.0-sglang-deepseek-v4-dev.1v1.2.0sglang-runtimeApr 25, 2026Superseded — recipe in v1.3.0Earliest DSv4 preview (SGLang, B200 only); superseded by dev.2/dev.3.yes / no / no / no

Platform-preview artifact coverage

PreviewImagesWheelsHelmCrates
v1.3.0-dev.1yesyesyesyes
v1.1.0-dev.3yesyesnono
v1.1.0-dev.2yesyesnono
v1.1.0-dev.1yesyesyesno

Platform support

  • GPU architectures: Blackwell, Hopper, Ada Lovelace, Ampere
  • OS: Ubuntu 24.04 (x86_64, ARM64) — Containers and wheels
  • OS: Ubuntu 22.04 (x86_64) — Wheels only
  • CSP: AWS — Amazon Linux 2023 (x86_64) — Containers and wheels
  • CPU architectures: x86_64, ARM64 (Ubuntu 24.04 only)
  • Wheels: Wheels are built in a manylinux_2_28 environment (AlmaLinux 8, glibc 2.28+) and validated on Ubuntu 22.04 and 24.04. They install on any Linux distribution with glibc 2.28+ (Debian 11+, RHEL 9, etc.), but only Ubuntu 22.04/24.04 are officially verified.

Release statistics

ReleasePRsContributorsFirst-time contributorsBreaking changesKnown issues
v1.4.0640127295119
v1.3.0930125242410
v1.2.060382-511
v1.1.0896113-820
v1.0.0-90344114
v0.9.0--14113
v0.8.0--20014
v0.7.0--207
v0.6.0--403

Nightlies

ai-dynamo and ai-dynamo-runtime nightly builds from main publish wheels tagged *.devYYYYMMDD (since Apr 24, 2026); kvbm joined the nightly train on Aug 2, 2026. Install with pip or uv using --pre and the NVIDIA extra-index pattern shown above. Runtime containers publish to the *-runtime-nightly repositories on NGC, under a dated YYYYMMDD-<shortsha> tag plus a rolling latest tag.

VersionDatePackagesNotes
1.5.0.dev20260831Aug 31, 2026ai-dynamo, ai-dynamo-runtime, kvbm-
1.5.0.dev20260830Aug 30, 2026ai-dynamo, ai-dynamo-runtime, kvbm-
1.5.0.dev20260829Aug 29, 2026ai-dynamo, ai-dynamo-runtime, kvbm-