Enterprise RAG Retrieval Scaling and Sizing Guide#
NVIDIA Enterprise Reference Architecture
- Abstract
- Enterprise RAG Retrieval - Accuracy, Performance and Cost Efficiency
- RAG Enterprise Challenge
- Build a scalable AI factory for RAG
- RAG Retrieval workflow
- Key Metrics for RAG Retrieval
- Target Use cases
- RAG Balanced Performance
- Key Benchmarks for RAG Retrieval
- Key Factors affecting Accuracy and Performance
- Deep Dive into Key Factors affecting Accuracy, Performance and Cost Efficiency
- Chunk Size Impact on Latency
- Model and Chunk Size Impact on Throughput (RPS)
- TopK and Reranker ON/OFF Impact
- RAG Scaling Guidelines
- RAG Sizing Guidelines
- RAG Benchmarking methodology
- RAG Retrieval Performance Results
- RAG Chat (128/128) Retrieval Performance
- RAG Chat Balanced Performance Scale 1X - 96X
- RAG Chat Time to First Token Latency
- RAG Chat Latency Constraints
- RAG Chat Inter-Token Latency
- RAG Chat E2E Request Latency
- RAG E2E Latency Degradation with Scale
- RAG Chat TCO and Scaling Efficiency
- RAG Summarization (1024/512) Performance
- Summary - Enterprise RAG Retrieval
- References