A Comprehensive Guide to NIM LLM Latency-Throughput Benchmarking# Benchmarking Guide Overview Executive Summary Introduction to LLM Inference Benchmarking Background on How LLM Inference Works Metrics Time to First Token End-to-End Request Latency Inter-token Latency Tokens Per Second Requests Per Second Parameters and Best Practices Use Cases Load Control Other Parameters Using AIPerf to Benchmark Set Up an OpenAI-Compatible Llama-3 Inference Service with NVIDIA NIM Set Up AIPerf and Warm Up: Benchmarking a Single Use Case Sweep through a Number of Use Cases Analyze the Output Interpret the Results Benchmarking LoRA Models Best Practices for Multi-LoRA Benchmarking