Performance and Memory for NVIDIA NeMo Retriever Reranking NIM#

To benchmark the performance of NVIDIA NeMo Retriever Reranking NIM, you can use the AIPerf tool.

To run a performance benchmark with representative input data, create a JSONL file where each line represents a ranking request with a query and one or more passages.

Example#

Create a file named rankings.jsonl that contains the following content.

{"texts": [{"name":"query","contents":["What was the first car ever driven?"]},{"name":"passages","contents":["The first car was driven in the late 19th century.","A steam-powered vehicle was demonstrated before gasoline automobiles became common."]}]}
{"texts": [{"name":"query","contents":["Who served as the 5th President of the United States of America?"]},{"name":"passages","contents":["James Monroe served as the fifth president of the United States.","George Washington was the first president of the United States."]}]}
{"texts": [{"name":"query","contents":["Is the Sydney Opera House located in Australia?"]},{"name":"passages","contents":["The Sydney Opera House is located in Sydney, Australia.","The Eiffel Tower is located in Paris, France."]}]}
{"texts": [{"name":"query","contents":["In what state did they film Shrek 2?"]},{"name":"passages","contents":["Principal voice recording for Shrek 2 took place in California.","DreamWorks Animation produced Shrek 2."]}]}

Install the verified AIPerf version in a Python environment that can reach the NIM endpoint.

pip install aiperf==0.9.0

Run the following command to benchmark the ranking endpoint by using AIPerf.

aiperf profile \
    --model nvidia/llama-nemotron-rerank-vl-1b-v2 \
    --endpoint-type nim_rankings \
    --url localhost:8000 \
    --input-file rankings.jsonl \
    --custom-dataset-type single_turn \
    --extra-inputs truncate:END \
    --request-count 20 \
    --concurrency 5

For the full set of command line options, refer to the AIPerf command line options documentation.

About the Measurements#

The following sections report the memory footprint and the benchmark results for each model. All latency measurements are reported in milliseconds.

The previous AIPerf example benchmarks ranking requests using custom queries and passages. The benchmark tables use a separate synthetic AIPerf benchmark matrix for consistent GPU comparisons. For text rows, batch size is text passages per request and throughput is text passages per second. For image rows, batch size is image passages per request, image size is the generated JPEG resolution, and throughput is image passages per second.

Llama Nemotron Rerank VL 1B v2#

Memory Footprint#

The following table contains the memory footprint data for Llama Nemotron Rerank VL 1B v2.

Compute Capability

Precision

Approximate GPU Memory Size (GiB)

12.0

fp8

6.09

10.0

fp8

6.21

9.0

fp8

6.12

8.0

fp16

5.67

NVIDIA RTX PRO 6000 Blackwell Server Edition#

The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA RTX PRO 6000 Blackwell Server Edition.

Precision

Input Type

Image Size

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

passage

514

20

1

85.4

85.3

85.8

86.2

230.0

FP8

passage

514

20

3

224.2

210.5

300.9

301.6

264.4

FP8

passage

514

20

5

370.7

379.5

420.4

524.7

264.0

FP8

passage

514

40

1

160.8

160.7

162.4

163.3

246.1

FP8

passage

514

40

3

445.9

450.1

453.0

454.0

265.4

FP8

passage

514

40

5

737.6

751.9

755.9

756.5

265.1

FP8

passage

515

10

1

40.5

40.5

40.8

40.8

236.8

FP8

passage

515

10

3

109.4

110.0

110.4

110.7

268.2

FP8

passage

515

10

5

185.9

187.8

188.6

189.2

263.3

FP8

image passage

336x336 JPEG

1

1

12.1

12.1

12.2

12.2

82.8

FP8

image passage

336x336 JPEG

1

3

26.0

26.1

26.3

26.3

38.7

FP8

image passage

336x336 JPEG

1

5

41.4

41.7

41.9

42.0

24.3

FP8

image passage

560x560 JPEG

1

1

12.1

12.1

12.1

12.1

82.9

FP8

image passage

560x560 JPEG

1

3

25.8

25.8

26.0

26.3

39.1

FP8

image passage

560x560 JPEG

1

5

41.1

41.2

41.5

41.8

24.8

FP8

image passage

1024x1024 JPEG

1

1

34.2

34.2

34.5

34.6

29.2

FP8

image passage

1024x1024 JPEG

1

3

88.6

89.2

90.4

90.5

11.3

FP8

image passage

1024x1024 JPEG

1

5

153.7

154.8

155.8

156.1

6.7

NVIDIA B200#

The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA B200.

Precision

Input Type

Image Size

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

passage

514

20

1

42.6

42.7

43.1

43.2

450.6

FP8

passage

514

20

3

90.1

83.0

120.9

121.7

648.8

FP8

passage

514

20

5

149.4

154.4

168.7

209.1

650.4

FP8

passage

514

40

1

83.2

83.1

84.2

84.3

470.5

FP8

passage

514

40

3

203.8

206.0

207.4

208.0

578.1

FP8

passage

514

40

5

336.8

343.6

346.1

346.4

579.0

FP8

passage

515

10

1

30.6

30.8

31.1

31.1

310.2

FP8

passage

515

10

3

56.3

56.6

56.9

57.0

512.6

FP8

passage

515

10

5

75.7

75.8

76.8

77.2

633.8

FP8

image passage

336x336 JPEG

1

1

8.0

8.0

8.2

8.2

124.5

FP8

image passage

336x336 JPEG

1

3

14.0

14.0

14.1

14.2

72.0

FP8

image passage

336x336 JPEG

1

5

21.9

21.8

22.1

22.6

46.3

FP8

image passage

560x560 JPEG

1

1

8.1

8.1

8.3

8.4

122.9

FP8

image passage

560x560 JPEG

1

3

13.7

13.7

14.0

14.1

73.3

FP8

image passage

560x560 JPEG

1

5

21.5

21.5

21.8

21.9

47.0

FP8

image passage

1024x1024 JPEG

1

1

19.4

19.2

20.3

20.3

51.6

FP8

image passage

1024x1024 JPEG

1

3

43.5

43.6

44.0

44.1

23.2

FP8

image passage

1024x1024 JPEG

1

5

74.8

75.1

76.1

76.8

13.7

NVIDIA H100 80GB HBM3#

The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA H100 80GB HBM3.

Precision

Input Type

Image Size

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

passage

514

20

1

64.9

64.9

65.5

65.6

300.1

FP8

passage

514

20

3

155.3

147.4

209.1

210.2

380.4

FP8

passage

514

20

5

255.3

263.8

287.9

357.0

382.4

FP8

passage

514

40

1

122.6

122.5

124.1

124.5

321.8

FP8

passage

514

40

3

318.2

321.3

323.7

324.9

371.3

FP8

passage

514

40

5

530.4

540.6

544.2

545.1

368.2

FP8

passage

515

10

1

37.0

37.0

37.5

37.6

259.9

FP8

passage

515

10

3

84.7

85.0

86.1

86.6

344.6

FP8

passage

515

10

5

128.8

130.0

131.4

131.8

378.6

FP8

image passage

336x336 JPEG

1

1

9.5

9.5

9.6

9.6

105.6

FP8

image passage

336x336 JPEG

1

3

18.1

18.0

18.2

18.5

55.6

FP8

image passage

336x336 JPEG

1

5

28.0

27.8

30.0

31.4

36.2

FP8

image passage

560x560 JPEG

1

1

9.5

9.5

9.6

9.7

104.9

FP8

image passage

560x560 JPEG

1

3

17.7

17.7

18.0

18.3

56.7

FP8

image passage

560x560 JPEG

1

5

27.3

27.4

27.7

27.8

37.1

FP8

image passage

1024x1024 JPEG

1

1

27.1

27.2

27.6

27.6

36.9

FP8

image passage

1024x1024 JPEG

1

3

69.5

69.3

69.8

70.5

14.5

FP8

image passage

1024x1024 JPEG

1

5

122.0

122.5

123.3

130.5

8.3

NVIDIA A100 SXM4 80GB#

The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA A100 SXM4 80GB.

Precision

Input Type

Image Size

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

passage

514

20

1

144.1

143.8

145.8

146.8

137.9

FP8

passage

514

20

3

377.8

355.3

508.7

510.0

157.5

FP8

passage

514

20

5

624.8

640.1

879.2

884.0

157.0

FP8

passage

514

40

1

275.8

275.8

279.2

280.3

144.4

FP8

passage

514

40

3

762.2

769.7

775.7

777.6

155.6

FP8

passage

514

40

5

1259.2

1284.2

1292.5

1296.4

155.5

FP8

passage

515

10

1

69.7

70.2

70.6

70.8

141.6

FP8

passage

515

10

3

191.6

192.9

193.9

194.1

154.7

FP8

passage

515

10

5

315.7

319.7

321.4

321.7

155.5

FP8

image passage

336x336 JPEG

1

1

14.7

14.6

15.0

15.0

67.9

FP8

image passage

336x336 JPEG

1

3

35.9

36.2

36.5

36.6

28.1

FP8

image passage

336x336 JPEG

1

5

58.5

59.0

59.6

59.7

17.3

FP8

image passage

560x560 JPEG

1

1

15.0

15.0

15.2

15.2

66.7

FP8

image passage

560x560 JPEG

1

3

35.6

35.7

36.4

36.4

28.3

FP8

image passage

560x560 JPEG

1

5

57.5

57.8

58.8

59.3

17.7

FP8

image passage

1024x1024 JPEG

1

1

57.3

57.3

58.2

58.3

17.4

FP8

image passage

1024x1024 JPEG

1

3

159.4

160.1

161.1

161.7

6.3

FP8

image passage

1024x1024 JPEG

1

5

273.8

275.8

277.3

277.5

3.8

NVIDIA L40S#

The following table contains the performance data for Llama Nemotron Rerank VL 1B v2 on NVIDIA L40S.

Precision

Input Type

Image Size

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

passage

514

20

1

143.9

143.7

145.3

146.2

137.3

FP8

passage

514

20

3

412.6

386.3

553.5

554.9

144.1

FP8

passage

514

20

5

689.2

707.0

783.5

798.3

142.3

FP8

passage

514

40

1

277.9

277.8

280.3

281.0

143.0

FP8

passage

514

40

3

805.0

812.4

818.4

820.1

147.2

FP8

passage

514

40

5

1331.8

1357.9

1366.4

1368.6

147.0

FP8

passage

515

10

1

64.6

64.6

65.6

65.9

151.3

FP8

passage

515

10

3

197.3

198.6

201.6

201.9

149.9

FP8

passage

515

10

5

342.1

346.9

350.1

351.5

143.7

FP8

image passage

336x336 JPEG

1

1

14.7

14.7

14.8

14.8

68.2

FP8

image passage

336x336 JPEG

1

3

35.8

35.9

36.4

36.5

28.1

FP8

image passage

336x336 JPEG

1

5

57.5

57.9

58.4

60.0

17.6

FP8

image passage

560x560 JPEG

1

1

14.9

14.9

15.0

15.0

67.3

FP8

image passage

560x560 JPEG

1

3

35.3

35.6

36.1

36.3

28.4

FP8

image passage

560x560 JPEG

1

5

58.2

58.6

58.9

59.1

17.5

FP8

image passage

1024x1024 JPEG

1

1

58.6

58.5

59.3

59.3

17.1

FP8

image passage

1024x1024 JPEG

1

3

173.5

174.8

175.4

175.6

5.8

FP8

image passage

1024x1024 JPEG

1

5

279.4

279.1

280.3

281.8

3.6

Llama Nemotron Rerank 1B v2#

Memory Footprint#

The following table contains the memory footprint data for Llama Nemotron Rerank 1B v2.

Compute Capability

Precision

Approximate GPU Memory Size (GiB)

12.0

fp8

3.68

12.0

fp16

7.59

10.0

fp8

3.91

10.0

fp16

6.69

9.0

fp8

3.65

9.0

fp16

6.51

8.9

fp8

3.56

8.9

fp16

5.84

8.6

fp16

6.06

8.0

fp16

6.53

NVIDIA H100 80GB HBM3#

The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA H100 80GB HBM3.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

512

10

1

33

31

37

37

307.3

FP8

512

10

3

62

63

72

74

474.9

FP8

512

10

5

103

104

113

116

471.7

FP8

512

20

1

57

59

64

65

351.1

FP8

512

20

3

124

123

139

140

475.7

FP8

512

20

5

206

207

227

230

477.7

FP8

512

40

1

99

99

109

110

402.8

FP8

512

40

3

244

248

261

267

483.4

FP8

512

40

5

405

414

429

434

483.7

FP16

512

10

1

40

39

43

44

250.4

FP16

512

10

3

88

89

96

104

337.7

FP16

512

10

5

145

148

153

154

337.1

FP16

512

20

1

71

74

78

79

280.1

FP16

512

20

3

171

172

184

187

345.0

FP16

512

20

5

284

288

305

310

345.3

FP16

512

40

1

129

129

139

140

309.5

FP16

512

40

3

341

346

358

361

347.5

FP16

512

40

5

565

575

592

601

347.2

NVIDIA A100 SXM4 80GB#

The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA A100 SXM4 80GB.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP16

512

10

1

74

74

79

80

134.8

FP16

512

10

3

185

188

202

244

157.3

FP16

512

10

5

311

315

325

329

157.6

FP16

512

20

1

139

137

149

150

143.5

FP16

512

20

3

371

373

394

398

159.8

FP16

512

20

5

615

622

644

648

159.5

FP16

512

40

1

267

266

286

290

149.8

FP16

512

40

3

744

752

788

799

159.1

FP16

512

40

5

1231

1252

1296

1316

159.0

NVIDIA A100 PCIe 80GB#

The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA A100 PCIe 80GB.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP16

512

10

1

79

79

84

85

126.4

FP16

512

10

3

203

206

217

271

144.5

FP16

512

10

5

339

346

354

358

144.4

FP16

512

20

1

151

151

160

161

132.0

FP16

512

20

3

405

411

428

435

145.4

FP16

512

20

5

672

685

710

717

145.5

FP16

512

40

1

289

287

307

310

138.2

FP16

512

40

3

811

817

842

851

146.3

FP16

512

40

5

1340

1357

1405

1412

146.2

NVIDIA L40S#

The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA L40S.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

512

10

1

58

58

61

61

170.8

FP8

512

10

3

149

150

159

204

197.8

FP8

512

10

5

246

250

257

260

198.8

FP8

512

20

1

119

117

123

124

168.5

FP8

512

20

3

315

319

325

326

187.9

FP8

512

20

5

523

532

540

540

187.5

FP8

512

40

1

234

234

242

243

171.1

FP8

512

40

3

652

661

670

674

181.6

FP8

512

40

5

1080

1101

1118

1120

181.4

FP16

512

10

1

78

78

80

80

128.2

FP16

512

10

3

205

207

210

212

144.8

FP16

512

10

5

340

346

354

357

144.0

FP16

512

20

1

158

157

162

163

126.8

FP16

512

20

3

430

436

443

446

137.2

FP16

512

20

5

716

728

739

744

136.9

FP16

512

40

1

312

312

320

321

128.0

FP16

512

40

3

886

896

907

910

134.0

FP16

512

40

5

1463

1492

1512

1515

134.0

FP16

512

40

5

1463

1492

1512

1515

134.0

NVIDIA L4#

The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA L4.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

512

10

1

190

187

197

199

52.7

FP8

512

10

3

534

543

554

557

55.2

FP8

512

10

5

888

907

924

925

55.2

FP8

512

20

1

383

384

395

397

52.1

FP8

512

20

3

1117

1134

1150

1153

53.0

FP8

512

20

5

1850

1885

1908

1916

53.0

FP8

512

40

1

768

768

785

787

52.1

FP8

512

40

3

2254

2279

2302

2308

52.7

FP8

512

40

5

3712

3795

3834

3839

52.7

FP16

512

10

1

188

188

196

197

53.2

FP16

512

10

3

533

541

554

559

55.4

FP16

512

10

5

884

902

917

922

55.4

FP16

512

20

1

383

382

396

399

52.2

FP16

512

20

3

1119

1134

1147

1152

53.1

FP16

512

20

5

1848

1884

1910

1916

53.0

FP16

512

40

1

767

767

784

788

52.2

FP16

512

40

3

2245

2276

2301

2308

52.8

FP16

512

40

5

3714

3790

3825

3831

52.8

NVIDIA A10G#

The following table contains the performance data for Llama Nemotron Rerank 1B v2 on NVIDIA A10G.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP16

512

10

1

230

229

240

244

43.5

FP16

512

10

3

639

644

667

675

46.5

FP16

512

10

5

1055

1076

1101

1107

46.4

FP16

512

20

1

447

442

464

468

44.7

FP16

512

20

3

1257

1276

1310

1320

46.9

FP16

512

20

5

2088

2126

2172

2185

46.9

FP16

512

40

1

877

879

902

906

45.6

FP16

512

40

3

2534

2566

2605

2617

46.7

FP16

512

40

5

4194

4273

4321

4339

46.7

Llama Nemotron Rerank 500m v2#

Memory Footprint#

The following table contains the memory footprint data for Llama Nemotron Rerank 500m v2.

GPU

Precision

Max Batch Size

Max Sequence Length

Approximate GPU Memory Size (GiB)

NVIDIA H100 80GB HBM3

fp8

1

8192

1.4

NVIDIA H100 80GB HBM3

fp8

8

8192

4.5

NVIDIA H100 80GB HBM3

fp8

16

8192

8.04

NVIDIA H100 80GB HBM3

fp8

30

1024

2.16

NVIDIA H100 80GB HBM3

fp8

30

2048

3.58

NVIDIA H100 80GB HBM3

fp8

30

4096

6.66

NVIDIA H100 80GB HBM3

fp8

30

8192

14.21

NVIDIA H100 80GB HBM3

fp16

1

8192

3.04

NVIDIA H100 80GB HBM3

fp16

8

8192

10.38

NVIDIA H100 80GB HBM3

fp16

16

8192

18.75

NVIDIA H100 80GB HBM3

fp16

30

1024

5.69

NVIDIA H100 80GB HBM3

fp16

30

2048

9.15

NVIDIA H100 80GB HBM3

fp16

30

4096

16.77

NVIDIA H100 80GB HBM3

fp16

30

8192

33.41

NVIDIA H100 NVL

fp8

1

8192

1.4

NVIDIA H100 NVL

fp8

8

8192

4.5

NVIDIA H100 NVL

fp8

16

8192

8.04

NVIDIA H100 NVL

fp8

30

1024

2.16

NVIDIA H100 NVL

fp8

30

2048

3.58

NVIDIA H100 NVL

fp8

30

4096

6.66

NVIDIA H100 NVL

fp8

30

8192

14.21

NVIDIA H100 NVL

fp16

1

8192

3.04

NVIDIA H100 NVL

fp16

8

8192

10.38

NVIDIA H100 NVL

fp16

16

8192

18.75

NVIDIA H100 NVL

fp16

30

1024

5.69

NVIDIA H100 NVL

fp16

30

2048

9.15

NVIDIA H100 NVL

fp16

30

4096

16.77

NVIDIA H100 NVL

fp16

30

8192

33.41

NVIDIA A100 SXM4 80GB

fp16

1

8192

2.53

NVIDIA A100 SXM4 80GB

fp16

8

8192

9.63

NVIDIA A100 SXM4 80GB

fp16

16

8192

18.25

NVIDIA A100 SXM4 80GB

fp16

30

1024

5.19

NVIDIA A100 SXM4 80GB

fp16

30

2048

8.65

NVIDIA A100 SXM4 80GB

fp16

30

4096

16.27

NVIDIA A100 SXM4 80GB

fp16

30

8192

32.91

NVIDIA A100 SXM4 40GB

fp16

1

8192

2.53

NVIDIA A100 SXM4 40GB

fp16

8

8192

9.63

NVIDIA A100 SXM4 40GB

fp16

16

8192

18.25

NVIDIA A100 SXM4 40GB

fp16

30

1024

5.19

NVIDIA A100 SXM4 40GB

fp16

30

2048

8.65

NVIDIA A100 SXM4 40GB

fp16

30

4096

16.27

NVIDIA A100 SXM4 40GB

fp16

30

8192

32.91

NVIDIA L40S

fp8

1

8192

1.52

NVIDIA L40S

fp8

8

8192

4.5

NVIDIA L40S

fp8

16

8192

8.03

NVIDIA L40S

fp8

30

1024

2.16

NVIDIA L40S

fp8

30

2048

3.58

NVIDIA L40S

fp8

30

4096

6.65

NVIDIA L40S

fp8

30

8192

14.21

NVIDIA L40S

fp16

1

8192

2.6

NVIDIA L40S

fp16

8

8192

9.69

NVIDIA L40S

fp16

16

8192

17.81

NVIDIA L40S

fp16

30

1024

5.14

NVIDIA L40S

fp16

30

2048

8.59

NVIDIA L40S

fp16

30

4096

15.86

NVIDIA L40S

fp16

30

8192

32.03

NVIDIA L4

fp16

1

8192

2.27

NVIDIA L4

fp16

8

8192

9.81

NVIDIA L4

fp16

16

1024

3.5

NVIDIA L4

fp16

16

2048

5.38

NVIDIA L4

fp16

16

4096

9.31

NVIDIA L4

fp16

16

8192

18.06

NVIDIA A10G

fp16

1

8192

2.78

NVIDIA A10G

fp16

8

8192

9.88

NVIDIA A10G

fp16

16

1024

3.66

NVIDIA A10G

fp16

16

2048

5.5

NVIDIA A10G

fp16

16

4096

9.38

NVIDIA A10G

fp16

16

8192

18.0

NVIDIA H100 80GB HBM3#

The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA H100 80GB HBM3.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

512

10

1

27.05

27

30

31.34

367

FP8

512

10

3

33.18

33

36

37.21

878

FP8

512

10

5

54.49

55

60

62.92

896

FP8

512

20

1

38.73

38

42

43.07

513

FP8

512

20

3

60.5

61

67

67.74

967

FP8

512

20

5

100.89

101

110

111.25

971

FP8

512

40

1

66.56

67

72

72.58

598

FP8

512

40

3

123.77

124

133

135.71

958

FP8

512

40

5

204.98

208

220

224.35

956

FP16

512

10

1

29.74

29

33

34.32

334

FP16

512

10

3

41.11

41

44

45.61

717

FP16

512

10

5

68.08

68

73

77.68

718

FP16

512

20

1

44.16

43

48

48.71

450

FP16

512

20

3

78.06

78

84

86.41

759

FP16

512

20

5

129.04

130

137

137.66

759

FP16

512

40

1

77.09

77

82

82.72

517

FP16

512

40

3

157.58

158

166

169.89

753

FP16

512

40

5

260.69

264

272

277.55

751

NVIDIA H100 NVL#

The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA H100 NVL.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

512

10

1

28.16

28

32

32.65

352

FP8

512

10

3

43.29

43

47

55.26

675

FP8

512

10

5

71.27

72

80

88.3

685

FP8

512

20

1

43.05

42

47

47.54

462

FP8

512

20

3

83.23

84

91

92.33

710

FP8

512

20

5

138.37

141

148

148.51

707

FP8

512

40

1

76.14

76

81

81.87

524

FP8

512

40

3

166.94

168

179

179.53

707

FP8

512

40

5

278.25

283

291

292.17

703

FP16

512

10

1

32.74

33

35

37.03

304

FP16

512

10

3

57.99

58

62

68.55

511

FP16

512

10

5

95.64

97

104

113.9

510

FP16

512

20

1

53.68

54

58

58.74

371

FP16

512

20

3

114.09

116

124

125.87

515

FP16

512

20

5

191.03

195

204

204.9

513

FP16

512

40

1

96.13

96

103

104.58

415

FP16

512

40

3

232.11

236

246

247.87

506

FP16

512

40

5

387.54

396

409

413.26

505

NVIDIA A100 SXM4 80GB#

The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA A100 SXM4 80GB.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP16

512

10

1

47.79

47

53

58.03

209

FP16

512

10

3

96.28

96

102

109.35

305

FP16

512

10

5

160.04

161

167

181.55

306

FP16

512

20

1

83.05

84

88

89.33

240

FP16

512

20

3

192.39

194

201

204.69

309

FP16

512

20

5

317.56

323

331

334.22

308

FP16

512

40

1

156.39

156

164

165.36

255

FP16

512

40

3

380.86

384

401

405.11

309

FP16

512

40

5

632.92

641

658

670.97

310

NVIDIA L40S#

The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA L40S.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP8

512

10

1

39.72

38

52

53.92

251

FP8

512

10

3

67.16

67

70

74.34

441

FP8

512

10

5

111.02

112

117

120.07

440

FP8

512

20

1

70.79

71

75

78.25

282

FP8

512

20

3

146.81

148

153

153.88

404

FP8

512

20

5

242.83

246

251

253.1

403

FP8

512

40

1

134.19

133

144

145.21

297

FP8

512

40

3

305.35

307

314

316.24

389

FP8

512

40

5

504.54

513

522

526.98

388

FP16

512

10

1

53.68

51

66

68.13

186

FP16

512

10

3

105.37

106

109

110.82

280

FP16

512

10

5

175.48

178

184

186.27

279

FP16

512

20

1

94.46

95

99

100.14

211

FP16

512

20

3

218.57

222

228

229.46

270

FP16

512

20

5

364.06

370

377

379.85

269

FP16

512

40

1

186.64

185

195

198.79

214

FP16

512

40

3

459.19

462

470

471.99

258

FP16

512

40

5

759.31

773

785

787.02

258

NVIDIA L4#

The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA L4.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP16

512

10

1

136.51

137

140

143.21

73

FP16

512

10

3

378.96

384

392

393.68

78

FP16

512

10

5

627.47

640

649

651.91

78

FP16

512

20

1

275.91

276

282

285.24

72

FP16

512

20

3

786.36

795

806

809.28

76

FP16

512

20

5

1298.23

1324

1339

1342.55

75

FP16

512

40

1

549.22

549

558

559.68

73

FP16

512

40

3

1567.23

1594

1609

1610.45

75

FP16

512

40

5

2603.41

2654

2677

2684.04

75

NVIDIA A10G#

The following table contains the performance data for Llama Nemotron Rerank 500m v2 on NVIDIA A10G.

Precision

Input Tokens

Batch Size

Concurrency

Avg Latency

P50 Latency

P90 Latency

P95 Latency

Throughput (inputs/s)

FP16

512

10

1

124.69

125

131

132.87

80

FP16

512

10

3

326.54

330

338

339.35

91

FP16

512

10

5

542.04

552

562

565.07

90

FP16

512

20

1

237.96

238

245

246.13

84

FP16

512

20

3

656.1

662

672

674.97

91

FP16

512

20

5

1084.72

1105

1119

1120.56

90

FP16

512

40

1

469.56

470

476

478.64

85

FP16

512

40

3

1308.15

1330

1345

1346.82

90

FP16

512

40

5

2173.49

2219

2235

2240.08

90