Performance and Memory for NVIDIA NeMo Retriever Reranking NIM#
To benchmark the performance of NVIDIA NeMo Retriever Reranking NIM, you can use the AIPerf tool.
To run a performance benchmark with representative input data, create a JSONL file where each line represents a ranking request with a query and one or more passages.
Example#
Create a file named rankings.jsonl that contains the following content.
{"texts": [{"name":"query","contents":["What was the first car ever driven?"]},{"name":"passages","contents":["The first car was driven in the late 19th century.","A steam-powered vehicle was demonstrated before gasoline automobiles became common."]}]}
{"texts": [{"name":"query","contents":["Who served as the 5th President of the United States of America?"]},{"name":"passages","contents":["James Monroe served as the fifth president of the United States.","George Washington was the first president of the United States."]}]}
{"texts": [{"name":"query","contents":["Is the Sydney Opera House located in Australia?"]},{"name":"passages","contents":["The Sydney Opera House is located in Sydney, Australia.","The Eiffel Tower is located in Paris, France."]}]}
{"texts": [{"name":"query","contents":["In what state did they film Shrek 2?"]},{"name":"passages","contents":["Principal voice recording for Shrek 2 took place in California.","DreamWorks Animation produced Shrek 2."]}]}
Install the verified AIPerf version in a Python environment that can reach the NIM endpoint.
pip install aiperf==0.9.0
Run the following command to benchmark the ranking endpoint by using AIPerf.
aiperf profile \
--model nvidia/llama-nemotron-rerank-vl-1b-v2 \
--endpoint-type nim_rankings \
--url localhost:8000 \
--input-file rankings.jsonl \
--custom-dataset-type single_turn \
--extra-inputs truncate:END \
--request-count 20 \
--concurrency 5
For the full set of command line options, refer to the AIPerf command line options documentation.
About the Measurements#
The following sections report the memory footprint and the benchmark results for each model. All latency measurements are reported in milliseconds.
The previous AIPerf example benchmarks ranking requests using custom queries and passages. The benchmark tables use a separate synthetic AIPerf benchmark matrix for consistent GPU comparisons. For text rows, batch size is text passages per request and throughput is text passages per second. For image rows, batch size is image passages per request, image size is the generated JPEG resolution, and throughput is image passages per second.
Llama Nemotron Rerank VL 1B v2#
Memory Footprint#
The following table contains the memory footprint data for Llama Nemotron Rerank VL 1B v2.
Precision |
Approximate GPU Memory Size (GiB) |
|
|---|---|---|
12.0 |
fp8 |
6.09 |
10.0 |
fp8 |
6.21 |
9.0 |
fp8 |
6.12 |
8.0 |
fp16 |
5.67 |
NVIDIA RTX PRO 6000 Blackwell Server Edition#
The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA RTX PRO 6000 Blackwell Server Edition.
Precision |
Input Type |
Image Size |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|---|---|
FP8 |
passage |
514 |
20 |
1 |
85.4 |
85.3 |
85.8 |
86.2 |
230.0 |
|
FP8 |
passage |
514 |
20 |
3 |
224.2 |
210.5 |
300.9 |
301.6 |
264.4 |
|
FP8 |
passage |
514 |
20 |
5 |
370.7 |
379.5 |
420.4 |
524.7 |
264.0 |
|
FP8 |
passage |
514 |
40 |
1 |
160.8 |
160.7 |
162.4 |
163.3 |
246.1 |
|
FP8 |
passage |
514 |
40 |
3 |
445.9 |
450.1 |
453.0 |
454.0 |
265.4 |
|
FP8 |
passage |
514 |
40 |
5 |
737.6 |
751.9 |
755.9 |
756.5 |
265.1 |
|
FP8 |
passage |
515 |
10 |
1 |
40.5 |
40.5 |
40.8 |
40.8 |
236.8 |
|
FP8 |
passage |
515 |
10 |
3 |
109.4 |
110.0 |
110.4 |
110.7 |
268.2 |
|
FP8 |
passage |
515 |
10 |
5 |
185.9 |
187.8 |
188.6 |
189.2 |
263.3 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
1 |
12.1 |
12.1 |
12.2 |
12.2 |
82.8 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
3 |
26.0 |
26.1 |
26.3 |
26.3 |
38.7 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
5 |
41.4 |
41.7 |
41.9 |
42.0 |
24.3 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
1 |
12.1 |
12.1 |
12.1 |
12.1 |
82.9 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
3 |
25.8 |
25.8 |
26.0 |
26.3 |
39.1 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
5 |
41.1 |
41.2 |
41.5 |
41.8 |
24.8 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
1 |
34.2 |
34.2 |
34.5 |
34.6 |
29.2 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
3 |
88.6 |
89.2 |
90.4 |
90.5 |
11.3 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
5 |
153.7 |
154.8 |
155.8 |
156.1 |
6.7 |
NVIDIA B200#
The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA B200.
Precision |
Input Type |
Image Size |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|---|---|
FP8 |
passage |
514 |
20 |
1 |
42.6 |
42.7 |
43.1 |
43.2 |
450.6 |
|
FP8 |
passage |
514 |
20 |
3 |
90.1 |
83.0 |
120.9 |
121.7 |
648.8 |
|
FP8 |
passage |
514 |
20 |
5 |
149.4 |
154.4 |
168.7 |
209.1 |
650.4 |
|
FP8 |
passage |
514 |
40 |
1 |
83.2 |
83.1 |
84.2 |
84.3 |
470.5 |
|
FP8 |
passage |
514 |
40 |
3 |
203.8 |
206.0 |
207.4 |
208.0 |
578.1 |
|
FP8 |
passage |
514 |
40 |
5 |
336.8 |
343.6 |
346.1 |
346.4 |
579.0 |
|
FP8 |
passage |
515 |
10 |
1 |
30.6 |
30.8 |
31.1 |
31.1 |
310.2 |
|
FP8 |
passage |
515 |
10 |
3 |
56.3 |
56.6 |
56.9 |
57.0 |
512.6 |
|
FP8 |
passage |
515 |
10 |
5 |
75.7 |
75.8 |
76.8 |
77.2 |
633.8 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
1 |
8.0 |
8.0 |
8.2 |
8.2 |
124.5 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
3 |
14.0 |
14.0 |
14.1 |
14.2 |
72.0 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
5 |
21.9 |
21.8 |
22.1 |
22.6 |
46.3 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
1 |
8.1 |
8.1 |
8.3 |
8.4 |
122.9 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
3 |
13.7 |
13.7 |
14.0 |
14.1 |
73.3 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
5 |
21.5 |
21.5 |
21.8 |
21.9 |
47.0 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
1 |
19.4 |
19.2 |
20.3 |
20.3 |
51.6 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
3 |
43.5 |
43.6 |
44.0 |
44.1 |
23.2 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
5 |
74.8 |
75.1 |
76.1 |
76.8 |
13.7 |
NVIDIA H100 80GB HBM3#
The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA H100 80GB HBM3.
Precision |
Input Type |
Image Size |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|---|---|
FP8 |
passage |
514 |
20 |
1 |
64.9 |
64.9 |
65.5 |
65.6 |
300.1 |
|
FP8 |
passage |
514 |
20 |
3 |
155.3 |
147.4 |
209.1 |
210.2 |
380.4 |
|
FP8 |
passage |
514 |
20 |
5 |
255.3 |
263.8 |
287.9 |
357.0 |
382.4 |
|
FP8 |
passage |
514 |
40 |
1 |
122.6 |
122.5 |
124.1 |
124.5 |
321.8 |
|
FP8 |
passage |
514 |
40 |
3 |
318.2 |
321.3 |
323.7 |
324.9 |
371.3 |
|
FP8 |
passage |
514 |
40 |
5 |
530.4 |
540.6 |
544.2 |
545.1 |
368.2 |
|
FP8 |
passage |
515 |
10 |
1 |
37.0 |
37.0 |
37.5 |
37.6 |
259.9 |
|
FP8 |
passage |
515 |
10 |
3 |
84.7 |
85.0 |
86.1 |
86.6 |
344.6 |
|
FP8 |
passage |
515 |
10 |
5 |
128.8 |
130.0 |
131.4 |
131.8 |
378.6 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
1 |
9.5 |
9.5 |
9.6 |
9.6 |
105.6 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
3 |
18.1 |
18.0 |
18.2 |
18.5 |
55.6 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
5 |
28.0 |
27.8 |
30.0 |
31.4 |
36.2 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
1 |
9.5 |
9.5 |
9.6 |
9.7 |
104.9 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
3 |
17.7 |
17.7 |
18.0 |
18.3 |
56.7 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
5 |
27.3 |
27.4 |
27.7 |
27.8 |
37.1 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
1 |
27.1 |
27.2 |
27.6 |
27.6 |
36.9 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
3 |
69.5 |
69.3 |
69.8 |
70.5 |
14.5 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
5 |
122.0 |
122.5 |
123.3 |
130.5 |
8.3 |
NVIDIA A100 SXM4 80GB#
The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA A100 SXM4 80GB.
Precision |
Input Type |
Image Size |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|---|---|
FP8 |
passage |
514 |
20 |
1 |
144.1 |
143.8 |
145.8 |
146.8 |
137.9 |
|
FP8 |
passage |
514 |
20 |
3 |
377.8 |
355.3 |
508.7 |
510.0 |
157.5 |
|
FP8 |
passage |
514 |
20 |
5 |
624.8 |
640.1 |
879.2 |
884.0 |
157.0 |
|
FP8 |
passage |
514 |
40 |
1 |
275.8 |
275.8 |
279.2 |
280.3 |
144.4 |
|
FP8 |
passage |
514 |
40 |
3 |
762.2 |
769.7 |
775.7 |
777.6 |
155.6 |
|
FP8 |
passage |
514 |
40 |
5 |
1259.2 |
1284.2 |
1292.5 |
1296.4 |
155.5 |
|
FP8 |
passage |
515 |
10 |
1 |
69.7 |
70.2 |
70.6 |
70.8 |
141.6 |
|
FP8 |
passage |
515 |
10 |
3 |
191.6 |
192.9 |
193.9 |
194.1 |
154.7 |
|
FP8 |
passage |
515 |
10 |
5 |
315.7 |
319.7 |
321.4 |
321.7 |
155.5 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
1 |
14.7 |
14.6 |
15.0 |
15.0 |
67.9 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
3 |
35.9 |
36.2 |
36.5 |
36.6 |
28.1 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
5 |
58.5 |
59.0 |
59.6 |
59.7 |
17.3 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
1 |
15.0 |
15.0 |
15.2 |
15.2 |
66.7 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
3 |
35.6 |
35.7 |
36.4 |
36.4 |
28.3 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
5 |
57.5 |
57.8 |
58.8 |
59.3 |
17.7 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
1 |
57.3 |
57.3 |
58.2 |
58.3 |
17.4 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
3 |
159.4 |
160.1 |
161.1 |
161.7 |
6.3 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
5 |
273.8 |
275.8 |
277.3 |
277.5 |
3.8 |
NVIDIA L40S#
The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA L40S.
Precision |
Input Type |
Image Size |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|---|---|
FP8 |
passage |
514 |
20 |
1 |
143.9 |
143.7 |
145.3 |
146.2 |
137.3 |
|
FP8 |
passage |
514 |
20 |
3 |
412.6 |
386.3 |
553.5 |
554.9 |
144.1 |
|
FP8 |
passage |
514 |
20 |
5 |
689.2 |
707.0 |
783.5 |
798.3 |
142.3 |
|
FP8 |
passage |
514 |
40 |
1 |
277.9 |
277.8 |
280.3 |
281.0 |
143.0 |
|
FP8 |
passage |
514 |
40 |
3 |
805.0 |
812.4 |
818.4 |
820.1 |
147.2 |
|
FP8 |
passage |
514 |
40 |
5 |
1331.8 |
1357.9 |
1366.4 |
1368.6 |
147.0 |
|
FP8 |
passage |
515 |
10 |
1 |
64.6 |
64.6 |
65.6 |
65.9 |
151.3 |
|
FP8 |
passage |
515 |
10 |
3 |
197.3 |
198.6 |
201.6 |
201.9 |
149.9 |
|
FP8 |
passage |
515 |
10 |
5 |
342.1 |
346.9 |
350.1 |
351.5 |
143.7 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
1 |
14.7 |
14.7 |
14.8 |
14.8 |
68.2 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
3 |
35.8 |
35.9 |
36.4 |
36.5 |
28.1 |
|
FP8 |
image passage |
336x336 JPEG |
1 |
5 |
57.5 |
57.9 |
58.4 |
60.0 |
17.6 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
1 |
14.9 |
14.9 |
15.0 |
15.0 |
67.3 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
3 |
35.3 |
35.6 |
36.1 |
36.3 |
28.4 |
|
FP8 |
image passage |
560x560 JPEG |
1 |
5 |
58.2 |
58.6 |
58.9 |
59.1 |
17.5 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
1 |
58.6 |
58.5 |
59.3 |
59.3 |
17.1 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
3 |
173.5 |
174.8 |
175.4 |
175.6 |
5.8 |
|
FP8 |
image passage |
1024x1024 JPEG |
1 |
5 |
279.4 |
279.1 |
280.3 |
281.8 |
3.6 |
Llama Nemotron Rerank 1B v2#
Memory Footprint#
The following table contains the memory footprint data for Llama Nemotron Rerank 1B v2.
Precision |
Approximate GPU Memory Size (GiB) |
|
|---|---|---|
12.0 |
fp8 |
3.68 |
12.0 |
fp16 |
7.59 |
10.0 |
fp8 |
3.91 |
10.0 |
fp16 |
6.69 |
9.0 |
fp8 |
3.65 |
9.0 |
fp16 |
6.51 |
8.9 |
fp8 |
3.56 |
8.9 |
fp16 |
5.84 |
8.6 |
fp16 |
6.06 |
8.0 |
fp16 |
6.53 |
NVIDIA H100 80GB HBM3#
The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA H100 80GB HBM3.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP8 |
512 |
10 |
1 |
33 |
31 |
37 |
37 |
307.3 |
FP8 |
512 |
10 |
3 |
62 |
63 |
72 |
74 |
474.9 |
FP8 |
512 |
10 |
5 |
103 |
104 |
113 |
116 |
471.7 |
FP8 |
512 |
20 |
1 |
57 |
59 |
64 |
65 |
351.1 |
FP8 |
512 |
20 |
3 |
124 |
123 |
139 |
140 |
475.7 |
FP8 |
512 |
20 |
5 |
206 |
207 |
227 |
230 |
477.7 |
FP8 |
512 |
40 |
1 |
99 |
99 |
109 |
110 |
402.8 |
FP8 |
512 |
40 |
3 |
244 |
248 |
261 |
267 |
483.4 |
FP8 |
512 |
40 |
5 |
405 |
414 |
429 |
434 |
483.7 |
FP16 |
512 |
10 |
1 |
40 |
39 |
43 |
44 |
250.4 |
FP16 |
512 |
10 |
3 |
88 |
89 |
96 |
104 |
337.7 |
FP16 |
512 |
10 |
5 |
145 |
148 |
153 |
154 |
337.1 |
FP16 |
512 |
20 |
1 |
71 |
74 |
78 |
79 |
280.1 |
FP16 |
512 |
20 |
3 |
171 |
172 |
184 |
187 |
345.0 |
FP16 |
512 |
20 |
5 |
284 |
288 |
305 |
310 |
345.3 |
FP16 |
512 |
40 |
1 |
129 |
129 |
139 |
140 |
309.5 |
FP16 |
512 |
40 |
3 |
341 |
346 |
358 |
361 |
347.5 |
FP16 |
512 |
40 |
5 |
565 |
575 |
592 |
601 |
347.2 |
NVIDIA A100 SXM4 80GB#
The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA A100 SXM4 80GB.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP16 |
512 |
10 |
1 |
74 |
74 |
79 |
80 |
134.8 |
FP16 |
512 |
10 |
3 |
185 |
188 |
202 |
244 |
157.3 |
FP16 |
512 |
10 |
5 |
311 |
315 |
325 |
329 |
157.6 |
FP16 |
512 |
20 |
1 |
139 |
137 |
149 |
150 |
143.5 |
FP16 |
512 |
20 |
3 |
371 |
373 |
394 |
398 |
159.8 |
FP16 |
512 |
20 |
5 |
615 |
622 |
644 |
648 |
159.5 |
FP16 |
512 |
40 |
1 |
267 |
266 |
286 |
290 |
149.8 |
FP16 |
512 |
40 |
3 |
744 |
752 |
788 |
799 |
159.1 |
FP16 |
512 |
40 |
5 |
1231 |
1252 |
1296 |
1316 |
159.0 |
NVIDIA A100 PCIe 80GB#
The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA A100 PCIe 80GB.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP16 |
512 |
10 |
1 |
79 |
79 |
84 |
85 |
126.4 |
FP16 |
512 |
10 |
3 |
203 |
206 |
217 |
271 |
144.5 |
FP16 |
512 |
10 |
5 |
339 |
346 |
354 |
358 |
144.4 |
FP16 |
512 |
20 |
1 |
151 |
151 |
160 |
161 |
132.0 |
FP16 |
512 |
20 |
3 |
405 |
411 |
428 |
435 |
145.4 |
FP16 |
512 |
20 |
5 |
672 |
685 |
710 |
717 |
145.5 |
FP16 |
512 |
40 |
1 |
289 |
287 |
307 |
310 |
138.2 |
FP16 |
512 |
40 |
3 |
811 |
817 |
842 |
851 |
146.3 |
FP16 |
512 |
40 |
5 |
1340 |
1357 |
1405 |
1412 |
146.2 |
NVIDIA L40S#
The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA L40S.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP8 |
512 |
10 |
1 |
58 |
58 |
61 |
61 |
170.8 |
FP8 |
512 |
10 |
3 |
149 |
150 |
159 |
204 |
197.8 |
FP8 |
512 |
10 |
5 |
246 |
250 |
257 |
260 |
198.8 |
FP8 |
512 |
20 |
1 |
119 |
117 |
123 |
124 |
168.5 |
FP8 |
512 |
20 |
3 |
315 |
319 |
325 |
326 |
187.9 |
FP8 |
512 |
20 |
5 |
523 |
532 |
540 |
540 |
187.5 |
FP8 |
512 |
40 |
1 |
234 |
234 |
242 |
243 |
171.1 |
FP8 |
512 |
40 |
3 |
652 |
661 |
670 |
674 |
181.6 |
FP8 |
512 |
40 |
5 |
1080 |
1101 |
1118 |
1120 |
181.4 |
FP16 |
512 |
10 |
1 |
78 |
78 |
80 |
80 |
128.2 |
FP16 |
512 |
10 |
3 |
205 |
207 |
210 |
212 |
144.8 |
FP16 |
512 |
10 |
5 |
340 |
346 |
354 |
357 |
144.0 |
FP16 |
512 |
20 |
1 |
158 |
157 |
162 |
163 |
126.8 |
FP16 |
512 |
20 |
3 |
430 |
436 |
443 |
446 |
137.2 |
FP16 |
512 |
20 |
5 |
716 |
728 |
739 |
744 |
136.9 |
FP16 |
512 |
40 |
1 |
312 |
312 |
320 |
321 |
128.0 |
FP16 |
512 |
40 |
3 |
886 |
896 |
907 |
910 |
134.0 |
FP16 |
512 |
40 |
5 |
1463 |
1492 |
1512 |
1515 |
134.0 |
FP16 |
512 |
40 |
5 |
1463 |
1492 |
1512 |
1515 |
134.0 |
NVIDIA L4#
The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA L4.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP8 |
512 |
10 |
1 |
190 |
187 |
197 |
199 |
52.7 |
FP8 |
512 |
10 |
3 |
534 |
543 |
554 |
557 |
55.2 |
FP8 |
512 |
10 |
5 |
888 |
907 |
924 |
925 |
55.2 |
FP8 |
512 |
20 |
1 |
383 |
384 |
395 |
397 |
52.1 |
FP8 |
512 |
20 |
3 |
1117 |
1134 |
1150 |
1153 |
53.0 |
FP8 |
512 |
20 |
5 |
1850 |
1885 |
1908 |
1916 |
53.0 |
FP8 |
512 |
40 |
1 |
768 |
768 |
785 |
787 |
52.1 |
FP8 |
512 |
40 |
3 |
2254 |
2279 |
2302 |
2308 |
52.7 |
FP8 |
512 |
40 |
5 |
3712 |
3795 |
3834 |
3839 |
52.7 |
FP16 |
512 |
10 |
1 |
188 |
188 |
196 |
197 |
53.2 |
FP16 |
512 |
10 |
3 |
533 |
541 |
554 |
559 |
55.4 |
FP16 |
512 |
10 |
5 |
884 |
902 |
917 |
922 |
55.4 |
FP16 |
512 |
20 |
1 |
383 |
382 |
396 |
399 |
52.2 |
FP16 |
512 |
20 |
3 |
1119 |
1134 |
1147 |
1152 |
53.1 |
FP16 |
512 |
20 |
5 |
1848 |
1884 |
1910 |
1916 |
53.0 |
FP16 |
512 |
40 |
1 |
767 |
767 |
784 |
788 |
52.2 |
FP16 |
512 |
40 |
3 |
2245 |
2276 |
2301 |
2308 |
52.8 |
FP16 |
512 |
40 |
5 |
3714 |
3790 |
3825 |
3831 |
52.8 |
NVIDIA A10G#
The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA A10G.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP16 |
512 |
10 |
1 |
230 |
229 |
240 |
244 |
43.5 |
FP16 |
512 |
10 |
3 |
639 |
644 |
667 |
675 |
46.5 |
FP16 |
512 |
10 |
5 |
1055 |
1076 |
1101 |
1107 |
46.4 |
FP16 |
512 |
20 |
1 |
447 |
442 |
464 |
468 |
44.7 |
FP16 |
512 |
20 |
3 |
1257 |
1276 |
1310 |
1320 |
46.9 |
FP16 |
512 |
20 |
5 |
2088 |
2126 |
2172 |
2185 |
46.9 |
FP16 |
512 |
40 |
1 |
877 |
879 |
902 |
906 |
45.6 |
FP16 |
512 |
40 |
3 |
2534 |
2566 |
2605 |
2617 |
46.7 |
FP16 |
512 |
40 |
5 |
4194 |
4273 |
4321 |
4339 |
46.7 |
Llama Nemotron Rerank 500m v2#
Memory Footprint#
The following table contains the memory footprint data for Llama Nemotron Rerank 500m v2.
GPU |
Precision |
Max Batch Size |
Max Sequence Length |
Approximate GPU Memory Size (GiB) |
|---|---|---|---|---|
NVIDIA H100 80GB HBM3 |
fp8 |
1 |
8192 |
1.4 |
NVIDIA H100 80GB HBM3 |
fp8 |
8 |
8192 |
4.5 |
NVIDIA H100 80GB HBM3 |
fp8 |
16 |
8192 |
8.04 |
NVIDIA H100 80GB HBM3 |
fp8 |
30 |
1024 |
2.16 |
NVIDIA H100 80GB HBM3 |
fp8 |
30 |
2048 |
3.58 |
NVIDIA H100 80GB HBM3 |
fp8 |
30 |
4096 |
6.66 |
NVIDIA H100 80GB HBM3 |
fp8 |
30 |
8192 |
14.21 |
NVIDIA H100 80GB HBM3 |
fp16 |
1 |
8192 |
3.04 |
NVIDIA H100 80GB HBM3 |
fp16 |
8 |
8192 |
10.38 |
NVIDIA H100 80GB HBM3 |
fp16 |
16 |
8192 |
18.75 |
NVIDIA H100 80GB HBM3 |
fp16 |
30 |
1024 |
5.69 |
NVIDIA H100 80GB HBM3 |
fp16 |
30 |
2048 |
9.15 |
NVIDIA H100 80GB HBM3 |
fp16 |
30 |
4096 |
16.77 |
NVIDIA H100 80GB HBM3 |
fp16 |
30 |
8192 |
33.41 |
NVIDIA H100 NVL |
fp8 |
1 |
8192 |
1.4 |
NVIDIA H100 NVL |
fp8 |
8 |
8192 |
4.5 |
NVIDIA H100 NVL |
fp8 |
16 |
8192 |
8.04 |
NVIDIA H100 NVL |
fp8 |
30 |
1024 |
2.16 |
NVIDIA H100 NVL |
fp8 |
30 |
2048 |
3.58 |
NVIDIA H100 NVL |
fp8 |
30 |
4096 |
6.66 |
NVIDIA H100 NVL |
fp8 |
30 |
8192 |
14.21 |
NVIDIA H100 NVL |
fp16 |
1 |
8192 |
3.04 |
NVIDIA H100 NVL |
fp16 |
8 |
8192 |
10.38 |
NVIDIA H100 NVL |
fp16 |
16 |
8192 |
18.75 |
NVIDIA H100 NVL |
fp16 |
30 |
1024 |
5.69 |
NVIDIA H100 NVL |
fp16 |
30 |
2048 |
9.15 |
NVIDIA H100 NVL |
fp16 |
30 |
4096 |
16.77 |
NVIDIA H100 NVL |
fp16 |
30 |
8192 |
33.41 |
NVIDIA A100 SXM4 80GB |
fp16 |
1 |
8192 |
2.53 |
NVIDIA A100 SXM4 80GB |
fp16 |
8 |
8192 |
9.63 |
NVIDIA A100 SXM4 80GB |
fp16 |
16 |
8192 |
18.25 |
NVIDIA A100 SXM4 80GB |
fp16 |
30 |
1024 |
5.19 |
NVIDIA A100 SXM4 80GB |
fp16 |
30 |
2048 |
8.65 |
NVIDIA A100 SXM4 80GB |
fp16 |
30 |
4096 |
16.27 |
NVIDIA A100 SXM4 80GB |
fp16 |
30 |
8192 |
32.91 |
NVIDIA A100 SXM4 40GB |
fp16 |
1 |
8192 |
2.53 |
NVIDIA A100 SXM4 40GB |
fp16 |
8 |
8192 |
9.63 |
NVIDIA A100 SXM4 40GB |
fp16 |
16 |
8192 |
18.25 |
NVIDIA A100 SXM4 40GB |
fp16 |
30 |
1024 |
5.19 |
NVIDIA A100 SXM4 40GB |
fp16 |
30 |
2048 |
8.65 |
NVIDIA A100 SXM4 40GB |
fp16 |
30 |
4096 |
16.27 |
NVIDIA A100 SXM4 40GB |
fp16 |
30 |
8192 |
32.91 |
NVIDIA L40S |
fp8 |
1 |
8192 |
1.52 |
NVIDIA L40S |
fp8 |
8 |
8192 |
4.5 |
NVIDIA L40S |
fp8 |
16 |
8192 |
8.03 |
NVIDIA L40S |
fp8 |
30 |
1024 |
2.16 |
NVIDIA L40S |
fp8 |
30 |
2048 |
3.58 |
NVIDIA L40S |
fp8 |
30 |
4096 |
6.65 |
NVIDIA L40S |
fp8 |
30 |
8192 |
14.21 |
NVIDIA L40S |
fp16 |
1 |
8192 |
2.6 |
NVIDIA L40S |
fp16 |
8 |
8192 |
9.69 |
NVIDIA L40S |
fp16 |
16 |
8192 |
17.81 |
NVIDIA L40S |
fp16 |
30 |
1024 |
5.14 |
NVIDIA L40S |
fp16 |
30 |
2048 |
8.59 |
NVIDIA L40S |
fp16 |
30 |
4096 |
15.86 |
NVIDIA L40S |
fp16 |
30 |
8192 |
32.03 |
NVIDIA L4 |
fp16 |
1 |
8192 |
2.27 |
NVIDIA L4 |
fp16 |
8 |
8192 |
9.81 |
NVIDIA L4 |
fp16 |
16 |
1024 |
3.5 |
NVIDIA L4 |
fp16 |
16 |
2048 |
5.38 |
NVIDIA L4 |
fp16 |
16 |
4096 |
9.31 |
NVIDIA L4 |
fp16 |
16 |
8192 |
18.06 |
NVIDIA A10G |
fp16 |
1 |
8192 |
2.78 |
NVIDIA A10G |
fp16 |
8 |
8192 |
9.88 |
NVIDIA A10G |
fp16 |
16 |
1024 |
3.66 |
NVIDIA A10G |
fp16 |
16 |
2048 |
5.5 |
NVIDIA A10G |
fp16 |
16 |
4096 |
9.38 |
NVIDIA A10G |
fp16 |
16 |
8192 |
18.0 |
NVIDIA H100 80GB HBM3#
The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA H100 80GB HBM3.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP8 |
512 |
10 |
1 |
27.05 |
27 |
30 |
31.34 |
367 |
FP8 |
512 |
10 |
3 |
33.18 |
33 |
36 |
37.21 |
878 |
FP8 |
512 |
10 |
5 |
54.49 |
55 |
60 |
62.92 |
896 |
FP8 |
512 |
20 |
1 |
38.73 |
38 |
42 |
43.07 |
513 |
FP8 |
512 |
20 |
3 |
60.5 |
61 |
67 |
67.74 |
967 |
FP8 |
512 |
20 |
5 |
100.89 |
101 |
110 |
111.25 |
971 |
FP8 |
512 |
40 |
1 |
66.56 |
67 |
72 |
72.58 |
598 |
FP8 |
512 |
40 |
3 |
123.77 |
124 |
133 |
135.71 |
958 |
FP8 |
512 |
40 |
5 |
204.98 |
208 |
220 |
224.35 |
956 |
FP16 |
512 |
10 |
1 |
29.74 |
29 |
33 |
34.32 |
334 |
FP16 |
512 |
10 |
3 |
41.11 |
41 |
44 |
45.61 |
717 |
FP16 |
512 |
10 |
5 |
68.08 |
68 |
73 |
77.68 |
718 |
FP16 |
512 |
20 |
1 |
44.16 |
43 |
48 |
48.71 |
450 |
FP16 |
512 |
20 |
3 |
78.06 |
78 |
84 |
86.41 |
759 |
FP16 |
512 |
20 |
5 |
129.04 |
130 |
137 |
137.66 |
759 |
FP16 |
512 |
40 |
1 |
77.09 |
77 |
82 |
82.72 |
517 |
FP16 |
512 |
40 |
3 |
157.58 |
158 |
166 |
169.89 |
753 |
FP16 |
512 |
40 |
5 |
260.69 |
264 |
272 |
277.55 |
751 |
NVIDIA H100 NVL#
The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA H100 NVL.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP8 |
512 |
10 |
1 |
28.16 |
28 |
32 |
32.65 |
352 |
FP8 |
512 |
10 |
3 |
43.29 |
43 |
47 |
55.26 |
675 |
FP8 |
512 |
10 |
5 |
71.27 |
72 |
80 |
88.3 |
685 |
FP8 |
512 |
20 |
1 |
43.05 |
42 |
47 |
47.54 |
462 |
FP8 |
512 |
20 |
3 |
83.23 |
84 |
91 |
92.33 |
710 |
FP8 |
512 |
20 |
5 |
138.37 |
141 |
148 |
148.51 |
707 |
FP8 |
512 |
40 |
1 |
76.14 |
76 |
81 |
81.87 |
524 |
FP8 |
512 |
40 |
3 |
166.94 |
168 |
179 |
179.53 |
707 |
FP8 |
512 |
40 |
5 |
278.25 |
283 |
291 |
292.17 |
703 |
FP16 |
512 |
10 |
1 |
32.74 |
33 |
35 |
37.03 |
304 |
FP16 |
512 |
10 |
3 |
57.99 |
58 |
62 |
68.55 |
511 |
FP16 |
512 |
10 |
5 |
95.64 |
97 |
104 |
113.9 |
510 |
FP16 |
512 |
20 |
1 |
53.68 |
54 |
58 |
58.74 |
371 |
FP16 |
512 |
20 |
3 |
114.09 |
116 |
124 |
125.87 |
515 |
FP16 |
512 |
20 |
5 |
191.03 |
195 |
204 |
204.9 |
513 |
FP16 |
512 |
40 |
1 |
96.13 |
96 |
103 |
104.58 |
415 |
FP16 |
512 |
40 |
3 |
232.11 |
236 |
246 |
247.87 |
506 |
FP16 |
512 |
40 |
5 |
387.54 |
396 |
409 |
413.26 |
505 |
NVIDIA A100 SXM4 80GB#
The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA A100 SXM4 80GB.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP16 |
512 |
10 |
1 |
47.79 |
47 |
53 |
58.03 |
209 |
FP16 |
512 |
10 |
3 |
96.28 |
96 |
102 |
109.35 |
305 |
FP16 |
512 |
10 |
5 |
160.04 |
161 |
167 |
181.55 |
306 |
FP16 |
512 |
20 |
1 |
83.05 |
84 |
88 |
89.33 |
240 |
FP16 |
512 |
20 |
3 |
192.39 |
194 |
201 |
204.69 |
309 |
FP16 |
512 |
20 |
5 |
317.56 |
323 |
331 |
334.22 |
308 |
FP16 |
512 |
40 |
1 |
156.39 |
156 |
164 |
165.36 |
255 |
FP16 |
512 |
40 |
3 |
380.86 |
384 |
401 |
405.11 |
309 |
FP16 |
512 |
40 |
5 |
632.92 |
641 |
658 |
670.97 |
310 |
NVIDIA L40S#
The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA L40S.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP8 |
512 |
10 |
1 |
39.72 |
38 |
52 |
53.92 |
251 |
FP8 |
512 |
10 |
3 |
67.16 |
67 |
70 |
74.34 |
441 |
FP8 |
512 |
10 |
5 |
111.02 |
112 |
117 |
120.07 |
440 |
FP8 |
512 |
20 |
1 |
70.79 |
71 |
75 |
78.25 |
282 |
FP8 |
512 |
20 |
3 |
146.81 |
148 |
153 |
153.88 |
404 |
FP8 |
512 |
20 |
5 |
242.83 |
246 |
251 |
253.1 |
403 |
FP8 |
512 |
40 |
1 |
134.19 |
133 |
144 |
145.21 |
297 |
FP8 |
512 |
40 |
3 |
305.35 |
307 |
314 |
316.24 |
389 |
FP8 |
512 |
40 |
5 |
504.54 |
513 |
522 |
526.98 |
388 |
FP16 |
512 |
10 |
1 |
53.68 |
51 |
66 |
68.13 |
186 |
FP16 |
512 |
10 |
3 |
105.37 |
106 |
109 |
110.82 |
280 |
FP16 |
512 |
10 |
5 |
175.48 |
178 |
184 |
186.27 |
279 |
FP16 |
512 |
20 |
1 |
94.46 |
95 |
99 |
100.14 |
211 |
FP16 |
512 |
20 |
3 |
218.57 |
222 |
228 |
229.46 |
270 |
FP16 |
512 |
20 |
5 |
364.06 |
370 |
377 |
379.85 |
269 |
FP16 |
512 |
40 |
1 |
186.64 |
185 |
195 |
198.79 |
214 |
FP16 |
512 |
40 |
3 |
459.19 |
462 |
470 |
471.99 |
258 |
FP16 |
512 |
40 |
5 |
759.31 |
773 |
785 |
787.02 |
258 |
NVIDIA L4#
The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA L4.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP16 |
512 |
10 |
1 |
136.51 |
137 |
140 |
143.21 |
73 |
FP16 |
512 |
10 |
3 |
378.96 |
384 |
392 |
393.68 |
78 |
FP16 |
512 |
10 |
5 |
627.47 |
640 |
649 |
651.91 |
78 |
FP16 |
512 |
20 |
1 |
275.91 |
276 |
282 |
285.24 |
72 |
FP16 |
512 |
20 |
3 |
786.36 |
795 |
806 |
809.28 |
76 |
FP16 |
512 |
20 |
5 |
1298.23 |
1324 |
1339 |
1342.55 |
75 |
FP16 |
512 |
40 |
1 |
549.22 |
549 |
558 |
559.68 |
73 |
FP16 |
512 |
40 |
3 |
1567.23 |
1594 |
1609 |
1610.45 |
75 |
FP16 |
512 |
40 |
5 |
2603.41 |
2654 |
2677 |
2684.04 |
75 |
NVIDIA A10G#
The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA A10G.
Precision |
Input Tokens |
Batch Size |
Concurrency |
Avg Latency |
P50 Latency |
P90 Latency |
P95 Latency |
Throughput (inputs/s) |
|---|---|---|---|---|---|---|---|---|
FP16 |
512 |
10 |
1 |
124.69 |
125 |
131 |
132.87 |
80 |
FP16 |
512 |
10 |
3 |
326.54 |
330 |
338 |
339.35 |
91 |
FP16 |
512 |
10 |
5 |
542.04 |
552 |
562 |
565.07 |
90 |
FP16 |
512 |
20 |
1 |
237.96 |
238 |
245 |
246.13 |
84 |
FP16 |
512 |
20 |
3 |
656.1 |
662 |
672 |
674.97 |
91 |
FP16 |
512 |
20 |
5 |
1084.72 |
1105 |
1119 |
1120.56 |
90 |
FP16 |
512 |
40 |
1 |
469.56 |
470 |
476 |
478.64 |
85 |
FP16 |
512 |
40 |
3 |
1308.15 |
1330 |
1345 |
1346.82 |
90 |
FP16 |
512 |
40 |
5 |
2173.49 |
2219 |
2235 |
2240.08 |
90 |