SGLang recipes for agentic traffic on B200, GB200, or H200, with FP8 KV caching, prefix caching, and speculative decoding.
SGLang1x B200 · 1x GB200 · 1x H200Agentic · MTP
Open recipe
Gemma-4-31B
Google / NVIDIA
Aggregated TensorRT-LLM recipes for long-context agentic traffic on B200, GB200, or H200, with KV-aware routing and CPU KV-cache offload.
TensorRT-LLM8x B200/GB200/H200Agentic + multimodal
Open recipevLLM recipes for agentic traffic on GB200 or H200, aggregated or 1P1D disaggregated with KV-aware routing, NIXL KV transfer, and H200 MTP7.
vLLM4x/8x GB200 · 8x/16x H200Agentic · up to 1M context
Open recipeDay-0 vLLM recipes for Moonshot’s Kimi-K3 on GB200 and GB300, aggregated or with prefill/decode disaggregation and KV-aware routing.
vLLMGB200 · GB300Day-0
Open recipevLLM and SGLang recipes for chat and agentic traffic on GB300 or GB200, aggregated or disaggregated, with KV-aware routing, MTP speculation, and FP8 weights/KV over MNNVL.
vLLM + SGLang16x GB300 · 16x GB200Chat + agentic · 262K context
Open recipeSGLang recipes for long-context agentic traffic on B200 (NVFP4) or H200 (FP8), aggregated or disaggregated, with KV-aware routing, EAGLE MTP speculation, and HiCache CPU offload.
SGLang4x/12x B200 · 8x/16x H200Agentic · up to 500K context
Open recipe
Inkling NVFP4
Thinking Machines
Thinking Machines’ first open-weights model — a multimodal MoE with controllable reasoning effort. vLLM GB300 targets serve agentic traffic at 1M context with MTP speculation and KV-aware routing, aggregated or disaggregated; the Day-0 SGLang B200 target adds image and audio input.
vLLM · SGLang8x GB300 / 8x B200Agentic · 1M contextText + image + audio
Open recipe
Kimi-K2.6
Moonshot / NVIDIA
Multi-target vLLM matrix covering B200 and H200 for chat and agentic traffic, with Eagle3 MLA speculation and KV-aware routing.
vLLM4x B200 / 8x H200Chat + agentic
Open recipe
Nemotron 3.5 Lightning
NVIDIA
30B hybrid Mamba/Attention/MoE recipe matrix for H100, H200, B200, and GB200, with NVFP4 and BF16 variants, vLLM MTP/DFlash/DSpark targets, DSpark with KV-routing, and experimental TensorRT-LLM MTP/no-spec.
vLLM + TRT-LLM1-4x GPUAgentic · 1M context
Open recipeAggregated vLLM targets for B200 and H200 chat and agentic traffic with MTP speculative decoding and trace-backed benchmarks.
vLLM4x B200 / 8x H200Chat + agentic
Open recipeNVFP4 and FP8 vLLM targets for B200 and H200 chat and agentic traffic with MTP speculative decoding and KV-aware routing.
vLLM4x B200 / 4x H200Chat + agentic
Open recipe20x GB200 SGLang P/D recipe for 1K input / 8K output traffic with EAGLE speculative decoding, plus an AWS EFA variant.
SGLang20x GB200Long output / static ISL-OSL
Open recipe
DeepSeek-V4-Pro
NVIDIA checkpoint / DeepSeek
vLLM agentic recipe — MoE 1.6T / 49B active, B200 (NVFP4, 1M ctx) and H200 (FP8), aggregated or disaggregated with MTP-2 and KV-aware routing.
vLLM8–32x B200/H200Agentic
Open recipe
DeepSeek-V4-Pro-0813
DeepSeek
vLLM agentic recipe — MoE 1.6T, MXFP4 experts + FP8 KV, GB200 and H200, aggregated or disaggregated with KV-aware routing. Full 1M context with no CPU KV offload.
vLLM8–16x GB200/H200Agentic1M ctx
Open recipe
DeepSeek-V4-Flash
NVIDIA checkpoint / DeepSeek
vLLM agentic recipe — MoE 284B / 13B active, B200 (NVFP4) and H200 (FP8), aggregated or disaggregated (2P1D / 4P3D) with KV-aware routing.
vLLM4–28x B200/H200Agentic
Open recipeTensorRT-LLM GB200 targets for static traffic plus vLLM B200/H200 agentic targets (agg + disagg) with KV-aware routing and EAGLE3 speculative decoding.
TRT-LLM + vLLMGB200 · B200 · H200Static + agentic
Open recipeDisaggregated vLLM serving with KV-aware routing on 16x H200, for multi-turn conversational traffic with prefix reuse.
vLLM16x H200Multi-turn conversationRelated benchmark
Open recipevLLM recipes for long-context agentic traffic on B200 NVFP4 and H200 FP8, aggregated or disaggregated 1P2D, with KV-aware routing and MTP speculation on H200 aggregated.
vLLM2-4x B200/H200Agentic trace replay
Open recipe16-GPU TensorRT-LLM recipe matrix for Hopper/Blackwell and aggregate/P-D serving.
TRT-LLM16x H100/H200Static ISL-OSL
Open recipeFP8 recipe set spanning TensorRT-LLM aggregate, TensorRT-LLM P/D, and vLLM P/D targets.
TRT-LLM2-8x GPUStatic ISL-OSL
Open recipeNo recipes match the selected filters.