Conclusion#
AI-Q blueprint is an agentic workflow that leverages an Enterprises existing multi modal data to create deep research reports along with internet search. The agentic workflow leverage the Nemotron Super 49B reasoning model to reflect and think about the results before generating a final report, this creates a very high volume of tokens per user session (Approx Average 19000 tokens). For deep research reports the time it takes to generate tokens is comparatively higher than most LLM chat/summary use cases. RTX PRO 6000 GPUs work for uptill peak usage of 100 users for a latency SLA of 1000 Seconds. Latency can grow for concurrencies above 128 users. To scale and get the most efficiency out of the System for AI-Q
Scaling the reasoning model, Nemotron 49B, has the biggest impact on lowering workflow TCO. As more users are added, scale the Nemotorn 49B NIM LLM pod to keep consistent workflow latency
Select NIM profiles with fp8 precision using TRT-LLM backend and Tensor Parallelism of 2 over vLLM backed profiles.
Linear scale can be achieved by scaling the reasoning LLM pods