Sizing Guidelines#

The above scaling methodology provides a framework on how to scale NIM LLMs on a system based on user response wait times under common input/output sequence length. Organizations can use this framework to run tests on their stack with workload characteristics that are more specific to their users and determine what CCU values they are getting for a single GPU, Fractional GPU and running multiple models simultaneously. Based on the results and expected peak load they can determine how many GPUs will be needed and correspondingly they can size the servers needed to hold that many GPUs. Organizations can also use this framework to determine scale and sizing for other LLM NIMs.

For example, For this version of the paper, below table would determine the concurrent users that run simultaneously using Run:ai on a single NVIDIA H100 NVL GPU, precision of fp8, characterization of 2000:200 on the 2-4-3-200 Enteprise RA based cluster and yet maintain a TTFT of close to 1000ms.

Table 8: Sizing GPUs for NIM Workloads with Run:ai

Model

CCU

Throughput (Output Tokens/Sec)

TTFT (ms)

Meta Llama 3.1 8B

137

2977

987

Deepseek R1 Distill Llama 8B

141

3225

987

Mixed (Both the above models)

278

6235

982