Abstract#

As enterprises scale their AI initiatives, maximizing return on investment from accelerated infrastructure has become a strategic and crucial objective. Most Enterprises are looking to run Inference services on Large Language models (LLMs) as a start but are also exploring running multiple models for varying use cases, to manage costs or to get better accuracy. Running multiple Inference services while using traditional methods of allocating GPUs statically often leads to underutilization, fragmented workloads, and increased operational overhead. This paper helps guide enterprises on how to pack more Inference models on a given set of NVIDIA GPUs using NVIDIA Run:ai, through intelligent scheduling, fractional GPUs, and dynamic resource management. We also explore the impact on performance with the Run:ai scheduler on utilizing fractional GPUs for NIM LLMs.