Benchmarking LLM Inference at Scale with AIPerf
At high concurrency, a benchmark client can saturate before the inference server does, and the numbers then describe the client rather than the deployment. In our blog post, Benchmarking LLM Inference at Scale with AIPerf, we introduce AIPerf, a ground-up rewrite and the successor to GenAI-Perf. AIPerf spreads load generation across worker processes and hands results to separate record processors, so the client keeps pace with the server. It supports more than 15 endpoint types, synthetic workloads, public datasets such as ShareGPT, and trace replay from Mooncake, Baseten, and WEKA AgentX. The post walks through a first benchmark against vLLM, explains how to read time to first token, inter-token latency, request latency, and output throughput, and then shapes traffic with Poisson arrivals and variable input and output lengths.