Enterprise RAG Ingestion Scaling Guide#
NVIDIA Enterprise Reference Architecture
- Abstract
- RAG Ingestion Pipeline Workflow
- RAG Ingestion Performance Benchmarking
- Benchmarking Methodology
- Basic Validation - RAG Ingestor-server (NV-Ingest) Text-only Data Ingestion
- Multimodal - NeMo Retriever (NV-Ingest) Extraction NIM Microservices
- Text-only - NeMo Retriever (NV-Ingest) Ingestion
- Wiki Text Bulk Ingest - Parquet Text-only Data
- RAG Ingestion Sizing
- Summary - Ingestion Best Practices
- Ingestion Pipeline - Time Distribution
- Text Chunking Strategy
- Embedding and VDB Best Practices
- Observability and Monitoring
- Multimodal Ingestion Best Practices
- GPU Sharing - Run:ai or MIG
- Scaling NV-Ingest and NIM Microservices
- Right Sizing NIM Fractions - Strong vs Weak Scaling
- Concurrency - Document (client) vs Worker (server)
- Data Preprocessing - Reduce E2E Ingestion Time
- Text-only NV-Ingest Best Practices
- Wiki Text Bulk Ingest Best Practices
- Ingestion Sizing - T-Shirt Sizing Guide
- References