Sizing Guideline#
Below is the sizing guidance for an Enterprise RA 2-8-5-200 Cluster with RTX PRO 6000 GPUs, a minimum of two nodes will allow for 16 simultaneous users to build reports while keeping total workflow latency below 1000 Seconds. As users grow, scaling up the Reasoning LLM pods is recommended to keep the workflow latency under 1000 seconds. This is based on the Biomedical and Finance datasets, as the datasets grow, certain components like Milvus, Nemotron reranker etc. will have to be scaled as well.
Table 4: Sizing RTX PRO 6000 BSE for AIQ on Enterprise RA 2-8-5-200 Architecture
Reasoning Model Scale |
Nodes (Worker) |
RTX PRO 6000 BSE GPUs |
Concurrency |
Workflow Latency (Seconds) |
LLM Latency (Seconds) |
Estimated Throughput (Cumulative Tokens) |
|---|---|---|---|---|---|---|
2X |
2 |
12 |
8 |
759.38 |
112.85 |
152000 |
4X |
2 |
16 |
16 |
869.99 |
117.38 |
304000 |
8X |
3 |
24 |
32 |
956.40 |
120 |
608000 |
16X |
5 |
40 |
64 |
951.35 |
116.60 |
1216000 |
32X |
9 |
72 |
100 |
1000 |
133 |
1900000 |