AI-Q Research Agent Blueprint# NVIDIA Enterprise Reference Architecture Abstract Introduction Scope Target Audience System Overview and Architecture AI-Q NVIDIA Research Agent Data Ingestion and Preparation Pipeline RAG (Retrieval-Augmented Generation) Framework AI-Q Instruct LLM and Toolchain Performance and Observability Enterprise Reference Architecture Overview (RA) Hardware Enterprise Reference Architecture Software Stack System Configuration Assumptions Pre-requisites for installing RAG and AI-Q Deploy and Configure RAG Blueprint Deploy and Configure AI-Q Blueprint Ingesting Enterprise Data for Research Benchmarking and Scale Methodology Setup NeMo Agent Toolkit to benchmark AI-Q Getting Started With Sizing a GPU Cluster Datasets needed for prompt during sizing Configuration File for NeMo Agent Toolkit to run sizing Run Benchmarking Scale Methodology Benchmarking and Scale testing Results for AI-Q Latency Impact of Reasoning Model Scale Latency v/s Concurrent Users at Scale AI-Q scales Linearly Latency Drops as Systems Scale Sizing Guideline Conclusion Appendix A Appendix B Appendix C Appendix D Notices Notices Notice Trademarks Copyright