vLLM recipes for agentic traffic on GB200 or H200, aggregated or 1P1D disaggregated with KV-aware routing, NIXL KV transfer, and H200 MTP7.
vLLM 4x/8x GB200 · 8x/16x H200 Agentic · up to 1M context
Open recipe Day-0 vLLM recipes for Moonshot’s Kimi-K3 on GB200 and GB300, aggregated or with prefill/decode disaggregation and KV-aware routing.
vLLM GB200 · GB300 Day-0
Open recipe vLLM and SGLang recipes for chat and agentic traffic on GB300 or GB200, aggregated or disaggregated, with KV-aware routing, MTP speculation, and FP8 weights/KV over MNNVL.
vLLM + SGLang 16x GB300 · 16x GB200 Chat + agentic · 262K context
Open recipe SGLang recipes for long-context agentic traffic on B200 (NVFP4) or H200 (FP8), aggregated or disaggregated, with KV-aware routing, EAGLE MTP speculation, and HiCache CPU offload.
SGLang 4x/12x B200 · 8x/16x H200 Agentic · up to 500K context
Open recipe Inkling NVFP4 Thinking Machines
Day-0 aggregated SGLang recipe for Thinking Machines’ first open-weights model — multimodal MoE with controllable reasoning effort, TP8 with EAGLE speculation.
SGLang 8x B200 Text + image + audio Day-0
Open recipe Kimi-K2.6 Moonshot / NVIDIA
Multi-target vLLM matrix covering B200 and H200 for chat and agentic traffic, with Eagle3 MLA speculation and KV-aware routing.
vLLM 4x B200 / 8x H200 Chat + agentic
Open recipe Nemotron 3.5 Lightning NVIDIA
30B hybrid Mamba/Attention/MoE recipe matrix for H100, H200, B200, and GB200, with NVFP4 and BF16 variants, vLLM MTP/DFlash/DSpark targets, DSpark with KV-routing, and experimental TensorRT-LLM MTP/no-spec.
vLLM + TRT-LLM 1-4x GPU Agentic · 1M context
Open recipe Aggregated vLLM targets for B200 and H200 chat and agentic traffic with MTP speculative decoding and trace-backed benchmarks.
vLLM 4x B200 / 8x H200 Chat + agentic
Open recipe NVFP4 and FP8 vLLM targets for B200 and H200 chat and agentic traffic with MTP speculative decoding and KV-aware routing.
vLLM 4x B200 / 4x H200 Chat + agentic
Open recipe 20x GB200 SGLang P/D recipe for 1K input / 8K output traffic with EAGLE speculative decoding, plus an AWS EFA variant.
SGLang 20x GB200 Long output / static ISL-OSL
Open recipe DeepSeek-V4-Pro NVIDIA checkpoint / DeepSeek
vLLM agentic recipe — MoE 1.6T / 49B active, B200 (NVFP4, 1M ctx) and H200 (FP8), aggregated or disaggregated with MTP-2 and KV-aware routing.
vLLM 8–32x B200/H200 Agentic
Open recipe DeepSeek-V4-Flash NVIDIA checkpoint / DeepSeek
vLLM agentic recipe — MoE 284B / 13B active, B200 (NVFP4) and H200 (FP8), aggregated or disaggregated (2P1D / 4P3D) with KV-aware routing.
vLLM 4–28x B200/H200 Agentic
Open recipe TensorRT-LLM GB200 targets for static traffic plus vLLM B200/H200 agentic targets (agg + disagg) with KV-aware routing and EAGLE3 speculative decoding.
TRT-LLM + vLLM GB200 · B200 · H200 Static + agentic
Open recipe Disaggregated vLLM serving with KV-aware routing on 16x H200, for multi-turn conversational traffic with prefix reuse.
vLLM 16x H200 Multi-turn conversation Related benchmark
Open recipe vLLM recipes for long-context agentic traffic on B200 NVFP4 and H200 FP8, aggregated or disaggregated 1P2D, with KV-aware routing and MTP speculation on H200 aggregated.
vLLM 2-4x B200/H200 Agentic trace replay
Open recipe 16-GPU TensorRT-LLM recipe matrix for Hopper/Blackwell and aggregate/P-D serving.
TRT-LLM 16x H100/H200 Static ISL-OSL
Open recipe FP8 recipe set spanning TensorRT-LLM aggregate, TensorRT-LLM P/D, and vLLM P/D targets.
TRT-LLM 2-8x GPU Static ISL-OSL
Open recipe No recipes match the selected filters.