Using AIPerf to Benchmark
NVIDIA AIPerf is a client-side generative AI benchmarking tool that reports TTFT, ITL, TPS, RPS, and related metrics. It works with any OpenAI-compatible inference service, including NVIDIA NIM. This section walks through benchmarking a Llama-3 model with AIPerf.
For metric definitions, refer to Metrics. For parameter guidance, refer to Parameters and Best Practices.
Set Up an OpenAI-Compatible Llama-3 Inference Service with NVIDIA NIM
NVIDIA NIM provides the easiest and quickest way to put an LLM into production. Refer to the NIM LLM getting started documentation for comprehensive guidance.
During startup, the NIM container downloads the required resources and begins serving the model behind an API endpoint. The following message indicates a successful startup:
After startup, NIM provides an OpenAI-compatible API that you can query, as shown in the following example:
NVIDIA benchmarking tests show that additional Docker flags can improve inference performance. Use either --security-opt seccomp=unconfined, which disables the Seccomp security profile, or --privileged, which grants the container broad host capabilities, including direct access to hardware, device files, and some kernel functionality. These flags improved performance by up to 5% with the NIM TensorRT-LLM v0.10.0 backend and up to 20% with the OSS vLLM backend tested on v0.4.3 or the NIM vLLM backend tested on NIM 1.0.0. NVIDIA verified this behavior on DGX A100 and H100 systems. These flags can reduce container security overhead, but they also elevate security risk. Use them only after reviewing your security requirements. Refer to the Docker privileged container documentation for more information.
Set Up AIPerf and Warm Up: Benchmarking a Single Use Case
After the NIM Llama-3 inference service is running, set up AIPerf. The easiest approach is a pre-built Docker container. Run AIPerf on the same host as NIM to avoid network latency, unless you intentionally want to include network effects in the measurement.
Refer to the AIPerf documentation for comprehensive setup guidance.
Run the following commands to use the NVIDIA Triton Server pre-built container.
After you start the container, install AIPerf:
This test uses the llama-3 tokenizer from Hugging Face, which is a guarded Meta-Llama-3.1-8B-Instruct repository. You need to apply for access, then log in with your Hugging Face credential.
Then, start the AIPerf evaluation harness, which runs a warm-up load test on the NIM backend:
This example specifies the input sequence length, output sequence length, and concurrency to test. It also tells the backend to ignore the EOS token so the output reaches the intended length.
Refer to the AIPerf CLI options documentation for the full set of options and parameters.
After successful execution, results similar to the following appear in the terminal:

Sweep through a Number of Use Cases
Benchmarking usually sweeps across use cases, such as input/output length combinations, and load scenarios, such as different concurrency values. Use the following Bash script to define the parameters so AIPerf runs all combinations.
Before running a full sweep, run a warm-up test as shown in the previous section.
Save this script in a working directory, such as /workdir/benchmark.sh. Run it with the following command:
The --request-count parameter specifies the number of requests for the measurement. This example sets it to three times the concurrency level to obtain a stable measurement. High concurrency on a large model with a large ISL/OSL scenario can significantly increase benchmark time.
Analyze the Output
When the tests complete, AIPerf writes structured outputs under an artifact directory in your mounted working directory (/workdir in these examples), organized by input/output length and concurrency. Your results should resemble the following.
The profile_export_aiperf.json files contain the main benchmarking results. Use the following Python snippet to load a result file for one use case:
You can also collect tokens-per-second and TTFT metrics across concurrencies for a given use case:
Finally, plot and analyze the latency-throughput curve with the collected data. Each point corresponds to a concurrency value.
The resulting plot from AIPerf measurement data looks like the following:

Interpret the Results
The previous plot shows TTFT on the x-axis, total system throughput on the y-axis, and concurrencies on each dot. You can use the plot in two ways:
- If you have a latency budget, use the maximum acceptable TTFT as the x value. The matching y value and concurrency show the highest throughput you can achieve within that latency limit.
- If you have a target concurrency, locate that dot on the graph. The matching x and y values show latency and throughput for that concurrency level.
The plot also shows concurrencies where latency grows quickly with little or no throughput gain. For example, in the plot above, concurrency=100 is one such value.
Similar plots can use ITL, e2e_latency, or TPS_per_user as the x-axis to show the trade-off between total system throughput and individual latency.