Recipes

Pick a model, choose the hardware/runtime target, and deploy the matching Dynamo recipe.

View as Markdown

Start with the model, runtime, and hardware you need to run. Model cards are sorted newest to oldest by release date. Use Feature Benchmarks when you want evidence that a Dynamo feature or topology helps under controlled traffic.

Recipe catalog

20 model families

Filter by provider, runtime, and hardware. Each card notes its serving technique and workload shape.

107Deployable configurations
Provider
Runtime
Hardware

Qwen3.6-35B-A3B

Qwen

SGLang recipes for agentic traffic on B200, GB200, or H200, with FP8 KV caching, prefix caching, and speculative decoding.

SGLang1x B200 · 1x GB200 · 1x H200Agentic · MTP
Open recipe

Gemma-4-31B

Google / NVIDIA

Aggregated TensorRT-LLM recipes for long-context agentic traffic on B200, GB200, or H200, with KV-aware routing and CPU KV-cache offload.

TensorRT-LLM8x B200/GB200/H200Agentic + multimodal
Open recipe

GLM-5.3-Flash

Z.AI

vLLM recipes for agentic traffic on GB200 or H200, aggregated or 1P1D disaggregated with KV-aware routing, NIXL KV transfer, and H200 MTP7.

vLLM4x/8x GB200 · 8x/16x H200Agentic · up to 1M context
Open recipe

Kimi-K3

Moonshot / NVIDIA

Day-0 vLLM recipes for Moonshot’s Kimi-K3 on GB200 and GB300, aggregated or with prefill/decode disaggregation and KV-aware routing.

vLLMGB200 · GB300Day-0
Open recipe

Qwen3.8-2.4T-A95B

Qwen

vLLM and SGLang recipes for chat and agentic traffic on GB300 or GB200, aggregated or disaggregated, with KV-aware routing, MTP speculation, and FP8 weights/KV over MNNVL.

vLLM + SGLang16x GB300 · 16x GB200Chat + agentic · 262K context
Open recipe

GLM-5.3/5.2

Z.AI / NVIDIA

SGLang recipes for long-context agentic traffic on B200 (NVFP4) or H200 (FP8), aggregated or disaggregated, with KV-aware routing, EAGLE MTP speculation, and HiCache CPU offload.

SGLang4x/12x B200 · 8x/16x H200Agentic · up to 500K context
Open recipe

Inkling NVFP4

Thinking Machines

Thinking Machines’ first open-weights model — a multimodal MoE with controllable reasoning effort. vLLM GB300 targets serve agentic traffic at 1M context with MTP speculation and KV-aware routing, aggregated or disaggregated; the Day-0 SGLang B200 target adds image and audio input.

vLLM · SGLang8x GB300 / 8x B200Agentic · 1M contextText + image + audio
Open recipe

Kimi-K2.6

Moonshot / NVIDIA

Multi-target vLLM matrix covering B200 and H200 for chat and agentic traffic, with Eagle3 MLA speculation and KV-aware routing.

vLLM4x B200 / 8x H200Chat + agentic
Open recipe

Nemotron 3.5 Lightning

NVIDIA

30B hybrid Mamba/Attention/MoE recipe matrix for H100, H200, B200, and GB200, with NVFP4 and BF16 variants, vLLM MTP/DFlash/DSpark targets, DSpark with KV-routing, and experimental TensorRT-LLM MTP/no-spec.

vLLM + TRT-LLM1-4x GPUAgentic · 1M context
Open recipe

Nemotron 3 Ultra

NVIDIA

Aggregated vLLM targets for B200 and H200 chat and agentic traffic with MTP speculative decoding and trace-backed benchmarks.

vLLM4x B200 / 8x H200Chat + agentic
Open recipe

Nemotron 3 Super

NVIDIA

NVFP4 and FP8 vLLM targets for B200 and H200 chat and agentic traffic with MTP speculative decoding and KV-aware routing.

vLLM4x B200 / 4x H200Chat + agentic
Open recipe

GLM-5 NVFP4

NVIDIA / Z.AI

20x GB200 SGLang P/D recipe for 1K input / 8K output traffic with EAGLE speculative decoding, plus an AWS EFA variant.

SGLang20x GB200Long output / static ISL-OSL
Open recipe

DeepSeek-V4-Pro

NVIDIA checkpoint / DeepSeek

vLLM agentic recipe — MoE 1.6T / 49B active, B200 (NVFP4, 1M ctx) and H200 (FP8), aggregated or disaggregated with MTP-2 and KV-aware routing.

vLLM8–32x B200/H200Agentic
Open recipe

DeepSeek-V4-Pro-0813

DeepSeek

vLLM agentic recipe — MoE 1.6T, MXFP4 experts + FP8 KV, GB200 and H200, aggregated or disaggregated with KV-aware routing. Full 1M context with no CPU KV offload.

vLLM8–16x GB200/H200Agentic1M ctx
Open recipe

DeepSeek-V4-Flash

NVIDIA checkpoint / DeepSeek

vLLM agentic recipe — MoE 284B / 13B active, B200 (NVFP4) and H200 (FP8), aggregated or disaggregated (2P1D / 4P3D) with KV-aware routing.

vLLM4–28x B200/H200Agentic
Open recipe

GPT-OSS-120B

OpenAI

TensorRT-LLM GB200 targets for static traffic plus vLLM B200/H200 agentic targets (agg + disagg) with KV-aware routing and EAGLE3 speculative decoding.

TRT-LLM + vLLMGB200 · B200 · H200Static + agentic
Open recipe

Qwen3-32B

Qwen

Disaggregated vLLM serving with KV-aware routing on 16x H200, for multi-turn conversational traffic with prefix reuse.

vLLM16x H200Multi-turn conversationRelated benchmark
Open recipe

Qwen3.5-122B-A10B

Qwen

vLLM recipes for long-context agentic traffic on B200 NVFP4 and H200 FP8, aggregated or disaggregated 1P2D, with KV-aware routing and MTP speculation on H200 aggregated.

vLLM2-4x B200/H200Agentic trace replay
Open recipe

Qwen3-235B-A22B FP8

Qwen

16-GPU TensorRT-LLM recipe matrix for Hopper/Blackwell and aggregate/P-D serving.

TRT-LLM16x H100/H200Static ISL-OSL
Open recipe

Qwen3-32B FP8

Qwen

FP8 recipe set spanning TensorRT-LLM aggregate, TensorRT-LLM P/D, and vLLM P/D targets.

TRT-LLM2-8x GPUStatic ISL-OSL
Open recipe

Use Feature Benchmarks when your question starts with a feature or performance claim, such as KV routing, embedding cache, speculative decoding, or topology. Recipe pages link back to Feature Benchmark pages whenever a recipe came from a comparison or feature-stack run.