Research publications
From the literature
Research publications
Papers that use, extend, or benchmark Dynamo, from the teams building on it and from the wider systems community.
SIGCOMM '26 NIACGeorgia TechAug 2026ARK: Avoiding Routing Collisions for KV Cache Transfer in Disaggregated LLM InferencearXivMarvellJul 29, 2026A Photonic-CXL Memory Appliance for Scalable KV Cache Management in LLM InferenceOSDI '26Korea UniversityJul 13, 2026Revisiting Pipeline Parallelism for LLM ServingarXivSolyx AIJun 13, 2026Solyx AI Grid: Hardware-Telemetry-Aware Routing Across Geographically Distributed GPU ClustersarXivJun 11, 2026The Price of Anarchy in Disaggregated InferenceMLSys 2026May 2026Breaking the Ice: Analyzing Cold Start Latency in vLLMMLSys 2026NVIDIAMay 2026A Pragmatic Exploration of Prefill-Decode Disaggregation in Large Scale InferenceFrontiers of Computer ScienceLenovoMay 9, 2026Adaptive Parallelism for LLM Inference with Model Irrelevant ProfilerEuroMLSys '26IBM ResearchApr 27, 2026A Case for a Simulation-Driven Exploration of Distributed GenAI PlatformsarXivApr 16, 2026Prefill-as-a-Service: KVCache of Next-Generation Models Could Go Cross-DatacenterarXivMar 13, 2026NCCL EP: Towards a Unified Expert Parallel Communication API for NCCLICML 2026Feb 16, 2026Efficient Multi-round LLM Inference over Disaggregated ServingarXivNVIDIAFeb 14, 2026ThunderAgent: A Simple, Fast and Program-Aware Agentic Inference SystemarXivNVIDIAJan 9, 2026AIConfigurator: Lightning-Fast Configuration Optimization for Multi-Framework LLM ServingIEEE AccessAmazon2026Optimizing GPU Workloads on Kubernetes: An Integrated Approach Using NVIDIA Dynamo, run:ai (KAI), and Amazon EKSarXivSK HynixDec 20, 2025TraCT: Disaggregated LLM Serving with CXL Shared Memory KV Cache at Rack-ScalearXivNov 6, 2025DuetServe: Harmonizing Prefill and Decode for LLM Serving via Adaptive GPU MultiplexingarXivCelestica AIOct 31, 2025AMD MI300X GPU Performance AnalysisarXivCapital OneOct 16, 2025From Attention to Disaggregation: Tracing the Evolution of LLM InferencearXivOct 15, 2025BanaServe: Unified KV Cache and Dynamic Module Migration for Balancing Disaggregated LLM Serving in AI InfrastructurearXivLMCacheOct 8, 2025LMCache: An Efficient KV Cache Layer for Enterprise-Scale LLM InferencearXivEPFLAug 22, 2025GreenLLM: SLO-Aware Dynamic Frequency Scaling for Energy-Efficient LLM ServingIEEE MicroMay 2025Toward Disaggregated and Heterogenous AI SystemsarXivInfinigence-AIApr 28, 2025semi-PD: Towards Efficient LLM Serving via Phase-Wise Disaggregated Computation and Unified Storage