Performance#
Version 2.6.0#
This section reports OpenFold2 v2.6.0 performance and accuracy, measured across all 18 supported SKUs with a performance and an accuracy run on each.
Note
The benchmark set changed in this release. Earlier sections use four protein chains (7WBN_A, 7ONG_A, 7ZHT_A, 7Y4I_A) measured by the previous benchmark client. v2.6.0 uses a 15-case set spanning 81 to 1596 residues, drawn from CASP15 and the PDB, and scored against experimental structures. Numbers in this section are therefore not directly comparable with those in the following sections, and no speedup-versus-2.5.0 column is given for that reason.
Benchmark Configuration#
Parameter |
Setting |
|---|---|
selected_models |
[1] |
use_templates |
false |
dataset |
15 monomers, 81-1596 residues |
metric |
|
Performance and Accuracy on H100#
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.46 |
0.773 |
T1104 |
115 |
5.07 |
0.856 |
T1106s1 |
120 |
5.16 |
0.758 |
7qsj-assembly1 |
373 |
20.08 |
0.935 |
T1137s1 |
408 |
23.42 |
0.803 |
T1112 |
459 |
25.19 |
0.836 |
T1114s3 |
534 |
34.57 |
0.889 |
7uww-assembly1 |
635 |
36.20 |
0.871 |
T1129s2 |
639 |
30.08 |
0.642 |
8k7x-assembly1 |
856 |
59.64 |
0.905 |
8t41-assembly1 |
867 |
67.44 |
0.920 |
8wnj-assembly1 |
920 |
76.31 |
0.897 |
T1125 |
1199 |
92.08 |
0.333 |
T1154 |
1423 |
127.60 |
0.366 |
8uxt-assembly1 |
1596 |
192.60 |
0.781 |
Measured on NVIDIA H100 80GB HBM3. predict_time is the end-to-end request
time for a single parameter set with no structural templates.
Note
lddt measures agreement with the experimental structure, and low values are a
property of the target rather than of the NIM: T1125 (0.333) and T1154 (0.366)
are CASP15 targets that are hard to predict from sequence. Read the accuracy
column per target, not as an average.
Performance Analysis#
Key Observations:
Scaling behavior: Runtime tracks sequence length across the measured range, from ~5 seconds at 81-120 residues to ~193 seconds at 1596 residues on H100.
Hardware range: Taking the H100 80GB HBM3 mean as 1.00x, GH200 is fastest at 0.81x, with GB300, GB200 and H200 at 0.90-0.92x and H200 NVL at 0.99x. B200 (1.04x), H100 NVL (1.11x) and B300 (1.17x) follow, then the RTX PRO 6000 Blackwell parts at 1.26-1.29x and H100 80GB PCIe at 1.42x. A100-class parts take 1.56-1.68x, L40S 1.83x and RTX 6000 Ada 2.11x. GB10 (DGX Spark) is 9.24x, which reflects its LPDDR5x unified memory rather than a configuration problem.
Short sequences: Below roughly 120 residues, runtime is close to flat (5.07-5.46 seconds on H100) because fixed per-request cost dominates the model itself.
Recommended Configuration:
For sequences beyond ~1200 residues, prefer 80GB-class or larger GPUs.
GB10 (DGX Spark) completes the full range including 1596 residues, but at substantially longer runtimes; size expectations accordingly.
Table 1: Performance and Accuracy Across the Supported NVIDIA Hardware Units#
One tab per SKU. Each shows all 15 benchmark cases sorted by sequence length,
with the end-to-end predict_time and the lddt scored against the
experimental structure. Mean is the average across the 15 cases.
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.46 |
0.773 |
T1104 |
115 |
5.07 |
0.856 |
T1106s1 |
120 |
5.16 |
0.758 |
7qsj-assembly1 |
373 |
20.08 |
0.935 |
T1137s1 |
408 |
23.42 |
0.803 |
T1112 |
459 |
25.19 |
0.836 |
T1114s3 |
534 |
34.57 |
0.889 |
7uww-assembly1 |
635 |
36.20 |
0.871 |
T1129s2 |
639 |
30.08 |
0.642 |
8k7x-assembly1 |
856 |
59.64 |
0.905 |
8t41-assembly1 |
867 |
67.44 |
0.920 |
8wnj-assembly1 |
920 |
76.31 |
0.897 |
T1125 |
1199 |
92.08 |
0.333 |
T1154 |
1423 |
127.60 |
0.366 |
8uxt-assembly1 |
1596 |
192.60 |
0.781 |
Mean |
— |
53.39 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.34 |
0.773 |
T1104 |
115 |
4.70 |
0.856 |
T1106s1 |
120 |
4.85 |
0.758 |
7qsj-assembly1 |
373 |
19.73 |
0.935 |
T1137s1 |
408 |
24.29 |
0.803 |
T1112 |
459 |
26.40 |
0.836 |
T1114s3 |
534 |
37.21 |
0.889 |
7uww-assembly1 |
635 |
39.76 |
0.871 |
T1129s2 |
639 |
32.19 |
0.642 |
8k7x-assembly1 |
856 |
66.30 |
0.905 |
8t41-assembly1 |
867 |
74.71 |
0.920 |
8wnj-assembly1 |
920 |
84.49 |
0.897 |
T1125 |
1199 |
103.17 |
0.333 |
T1154 |
1423 |
144.04 |
0.366 |
8uxt-assembly1 |
1596 |
220.16 |
0.781 |
Mean |
— |
59.16 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.50 |
0.773 |
T1104 |
115 |
4.79 |
0.850 |
T1106s1 |
120 |
4.96 |
0.758 |
7qsj-assembly1 |
373 |
23.39 |
0.935 |
T1137s1 |
408 |
28.30 |
0.803 |
T1112 |
459 |
31.45 |
0.836 |
T1114s3 |
534 |
44.83 |
0.889 |
7uww-assembly1 |
635 |
49.02 |
0.872 |
T1129s2 |
639 |
41.73 |
0.638 |
8k7x-assembly1 |
856 |
83.55 |
0.905 |
8t41-assembly1 |
867 |
94.31 |
0.920 |
8wnj-assembly1 |
920 |
105.91 |
0.897 |
T1125 |
1199 |
137.25 |
0.333 |
T1154 |
1423 |
195.73 |
0.366 |
8uxt-assembly1 |
1596 |
283.67 |
0.782 |
Mean |
— |
75.63 |
0.770 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
4.55 |
0.773 |
T1104 |
115 |
4.17 |
0.856 |
T1106s1 |
120 |
4.29 |
0.758 |
7qsj-assembly1 |
373 |
17.41 |
0.935 |
T1137s1 |
408 |
21.30 |
0.803 |
T1112 |
459 |
22.85 |
0.836 |
T1114s3 |
534 |
31.77 |
0.889 |
7uww-assembly1 |
635 |
33.62 |
0.871 |
T1129s2 |
639 |
27.63 |
0.642 |
8k7x-assembly1 |
856 |
55.06 |
0.905 |
8t41-assembly1 |
867 |
63.94 |
0.920 |
8wnj-assembly1 |
920 |
71.05 |
0.897 |
T1125 |
1199 |
84.36 |
0.333 |
T1154 |
1423 |
118.36 |
0.366 |
8uxt-assembly1 |
1596 |
177.79 |
0.781 |
Mean |
— |
49.21 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
4.63 |
0.773 |
T1104 |
115 |
4.18 |
0.856 |
T1106s1 |
120 |
4.32 |
0.758 |
7qsj-assembly1 |
373 |
18.89 |
0.935 |
T1137s1 |
408 |
22.46 |
0.803 |
T1112 |
459 |
24.30 |
0.836 |
T1114s3 |
534 |
34.81 |
0.889 |
7uww-assembly1 |
635 |
35.56 |
0.871 |
T1129s2 |
639 |
29.23 |
0.642 |
8k7x-assembly1 |
856 |
58.37 |
0.905 |
8t41-assembly1 |
867 |
66.81 |
0.920 |
8wnj-assembly1 |
920 |
75.27 |
0.897 |
T1125 |
1199 |
90.53 |
0.333 |
T1154 |
1423 |
127.69 |
0.366 |
8uxt-assembly1 |
1596 |
192.13 |
0.781 |
Mean |
— |
52.61 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.58 |
0.755 |
T1104 |
115 |
5.45 |
0.851 |
T1106s1 |
120 |
5.67 |
0.760 |
7qsj-assembly1 |
373 |
18.61 |
0.935 |
T1137s1 |
408 |
21.67 |
0.804 |
T1112 |
459 |
24.87 |
0.836 |
T1114s3 |
534 |
32.97 |
0.889 |
7uww-assembly1 |
635 |
36.82 |
0.872 |
T1129s2 |
639 |
33.39 |
0.657 |
8k7x-assembly1 |
856 |
53.97 |
0.900 |
8t41-assembly1 |
867 |
67.42 |
0.920 |
8wnj-assembly1 |
920 |
67.07 |
0.897 |
T1125 |
1199 |
105.92 |
0.331 |
T1154 |
1423 |
150.81 |
0.373 |
8uxt-assembly1 |
1596 |
203.94 |
0.784 |
Mean |
— |
55.61 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.76 |
0.755 |
T1104 |
115 |
5.58 |
0.851 |
T1106s1 |
120 |
5.98 |
0.760 |
7qsj-assembly1 |
373 |
18.73 |
0.935 |
T1137s1 |
408 |
26.77 |
0.804 |
T1112 |
459 |
29.95 |
0.836 |
T1114s3 |
534 |
36.29 |
0.889 |
7uww-assembly1 |
635 |
41.09 |
0.872 |
T1129s2 |
639 |
37.58 |
0.657 |
8k7x-assembly1 |
856 |
64.88 |
0.900 |
8t41-assembly1 |
867 |
84.29 |
0.920 |
8wnj-assembly1 |
920 |
82.14 |
0.897 |
T1125 |
1199 |
112.71 |
0.331 |
T1154 |
1423 |
151.14 |
0.373 |
8uxt-assembly1 |
1596 |
235.58 |
0.784 |
Mean |
— |
62.56 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
8.59 |
0.755 |
T1104 |
115 |
6.87 |
0.851 |
T1106s1 |
120 |
7.56 |
0.760 |
7qsj-assembly1 |
373 |
16.38 |
0.936 |
T1137s1 |
408 |
18.68 |
0.804 |
T1112 |
459 |
21.53 |
0.836 |
T1114s3 |
534 |
27.43 |
0.889 |
7uww-assembly1 |
635 |
31.71 |
0.872 |
T1129s2 |
639 |
28.77 |
0.651 |
8k7x-assembly1 |
856 |
45.30 |
0.900 |
8t41-assembly1 |
867 |
56.12 |
0.920 |
8wnj-assembly1 |
920 |
53.95 |
0.897 |
T1125 |
1199 |
93.53 |
0.332 |
T1154 |
1423 |
133.95 |
0.373 |
8uxt-assembly1 |
1596 |
173.55 |
0.784 |
Mean |
— |
48.26 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
7.55 |
0.755 |
T1104 |
115 |
6.55 |
0.851 |
T1106s1 |
120 |
6.83 |
0.760 |
7qsj-assembly1 |
373 |
16.36 |
0.936 |
T1137s1 |
408 |
18.55 |
0.804 |
T1112 |
459 |
21.47 |
0.836 |
T1114s3 |
534 |
26.99 |
0.889 |
7uww-assembly1 |
635 |
31.55 |
0.872 |
T1129s2 |
639 |
28.60 |
0.651 |
8k7x-assembly1 |
856 |
44.76 |
0.900 |
8t41-assembly1 |
867 |
55.66 |
0.920 |
8wnj-assembly1 |
920 |
53.10 |
0.897 |
T1125 |
1199 |
93.30 |
0.332 |
T1154 |
1423 |
133.42 |
0.373 |
8uxt-assembly1 |
1596 |
172.76 |
0.784 |
Mean |
— |
47.83 |
0.771 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
6.27 |
0.773 |
T1104 |
115 |
5.07 |
0.856 |
T1106s1 |
120 |
5.13 |
0.758 |
7qsj-assembly1 |
373 |
15.89 |
0.935 |
T1137s1 |
408 |
18.16 |
0.803 |
T1112 |
459 |
19.56 |
0.836 |
T1114s3 |
534 |
26.45 |
0.889 |
7uww-assembly1 |
635 |
29.11 |
0.871 |
T1129s2 |
639 |
24.47 |
0.635 |
8k7x-assembly1 |
856 |
47.49 |
0.905 |
8t41-assembly1 |
867 |
53.04 |
0.920 |
8wnj-assembly1 |
920 |
57.90 |
0.897 |
T1125 |
1199 |
75.90 |
0.333 |
T1154 |
1423 |
108.01 |
0.366 |
8uxt-assembly1 |
1596 |
153.06 |
0.780 |
Mean |
— |
43.03 |
0.770 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
8.57 |
0.772 |
T1104 |
115 |
6.75 |
0.844 |
T1106s1 |
120 |
6.87 |
0.761 |
7qsj-assembly1 |
373 |
30.96 |
0.936 |
T1137s1 |
408 |
35.10 |
0.797 |
T1112 |
459 |
38.40 |
0.837 |
T1114s3 |
534 |
53.28 |
0.889 |
7uww-assembly1 |
635 |
57.13 |
0.871 |
T1129s2 |
639 |
48.00 |
0.634 |
8k7x-assembly1 |
856 |
93.79 |
0.900 |
8t41-assembly1 |
867 |
105.45 |
0.920 |
8wnj-assembly1 |
920 |
117.11 |
0.897 |
T1125 |
1199 |
143.70 |
0.333 |
T1154 |
1423 |
203.14 |
0.349 |
8uxt-assembly1 |
1596 |
302.82 |
0.780 |
Mean |
— |
83.40 |
0.768 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
6.24 |
0.772 |
T1104 |
115 |
6.15 |
0.844 |
T1106s1 |
120 |
6.33 |
0.761 |
7qsj-assembly1 |
373 |
29.62 |
0.936 |
T1137s1 |
408 |
33.73 |
0.797 |
T1112 |
459 |
37.96 |
0.837 |
T1114s3 |
534 |
52.23 |
0.889 |
7uww-assembly1 |
635 |
57.85 |
0.871 |
T1129s2 |
639 |
49.51 |
0.634 |
8k7x-assembly1 |
856 |
96.09 |
0.900 |
8t41-assembly1 |
867 |
108.32 |
0.920 |
8wnj-assembly1 |
920 |
119.63 |
0.897 |
T1125 |
1199 |
152.34 |
0.333 |
T1154 |
1423 |
216.01 |
0.349 |
8uxt-assembly1 |
1596 |
312.46 |
0.780 |
Mean |
— |
85.63 |
0.768 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.74 |
0.772 |
T1104 |
115 |
5.96 |
0.844 |
T1106s1 |
120 |
6.30 |
0.761 |
7qsj-assembly1 |
373 |
30.97 |
0.936 |
T1137s1 |
408 |
34.76 |
0.797 |
T1112 |
459 |
39.67 |
0.837 |
T1114s3 |
534 |
53.73 |
0.889 |
7uww-assembly1 |
635 |
61.08 |
0.871 |
T1129s2 |
639 |
51.95 |
0.634 |
8k7x-assembly1 |
856 |
101.63 |
0.900 |
8t41-assembly1 |
867 |
113.16 |
0.920 |
8wnj-assembly1 |
920 |
124.81 |
0.897 |
T1125 |
1199 |
160.24 |
0.333 |
T1154 |
1423 |
226.20 |
0.349 |
8uxt-assembly1 |
1596 |
327.66 |
0.780 |
Mean |
— |
89.59 |
0.768 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
5.47 |
0.755 |
T1104 |
115 |
6.03 |
0.855 |
T1106s1 |
120 |
6.37 |
0.760 |
7qsj-assembly1 |
373 |
32.15 |
0.930 |
T1137s1 |
408 |
37.37 |
0.797 |
T1112 |
459 |
43.14 |
0.836 |
T1114s3 |
534 |
57.79 |
0.889 |
7uww-assembly1 |
635 |
68.17 |
0.872 |
T1129s2 |
639 |
61.94 |
0.646 |
8k7x-assembly1 |
856 |
113.00 |
0.898 |
8t41-assembly1 |
867 |
120.63 |
0.920 |
8wnj-assembly1 |
920 |
135.59 |
0.897 |
T1125 |
1199 |
179.13 |
0.334 |
T1154 |
1423 |
246.99 |
0.307 |
8uxt-assembly1 |
1596 |
349.03 |
0.783 |
Mean |
— |
97.52 |
0.765 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
6.14 |
0.773 |
T1104 |
115 |
5.71 |
0.850 |
T1106s1 |
120 |
5.65 |
0.760 |
7qsj-assembly1 |
373 |
22.70 |
0.935 |
T1137s1 |
408 |
25.94 |
0.797 |
T1112 |
459 |
29.75 |
0.836 |
T1114s3 |
534 |
40.00 |
0.889 |
7uww-assembly1 |
635 |
46.85 |
0.871 |
T1129s2 |
639 |
43.27 |
0.639 |
8k7x-assembly1 |
856 |
73.73 |
0.905 |
8t41-assembly1 |
867 |
85.09 |
0.920 |
8wnj-assembly1 |
920 |
90.73 |
0.897 |
T1125 |
1199 |
129.33 |
0.324 |
T1154 |
1423 |
180.91 |
0.342 |
8uxt-assembly1 |
1596 |
243.94 |
0.782 |
Mean |
— |
68.65 |
0.768 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
4.59 |
0.773 |
T1104 |
115 |
4.03 |
0.850 |
T1106s1 |
120 |
4.26 |
0.760 |
7qsj-assembly1 |
373 |
20.26 |
0.935 |
T1137s1 |
408 |
23.91 |
0.797 |
T1112 |
459 |
27.27 |
0.836 |
T1114s3 |
534 |
39.10 |
0.889 |
7uww-assembly1 |
635 |
44.03 |
0.871 |
T1129s2 |
639 |
40.35 |
0.639 |
8k7x-assembly1 |
856 |
72.66 |
0.905 |
8t41-assembly1 |
867 |
83.87 |
0.920 |
8wnj-assembly1 |
920 |
90.64 |
0.897 |
T1125 |
1199 |
126.31 |
0.324 |
T1154 |
1423 |
178.72 |
0.342 |
8uxt-assembly1 |
1596 |
247.49 |
0.782 |
Mean |
— |
67.17 |
0.768 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
6.20 |
0.772 |
T1104 |
115 |
5.71 |
0.853 |
T1106s1 |
120 |
5.98 |
0.747 |
7qsj-assembly1 |
373 |
34.29 |
0.936 |
T1137s1 |
408 |
41.51 |
0.796 |
T1112 |
459 |
47.51 |
0.836 |
T1114s3 |
534 |
69.12 |
0.889 |
7uww-assembly1 |
635 |
75.13 |
0.871 |
T1129s2 |
639 |
61.98 |
0.636 |
8k7x-assembly1 |
856 |
131.05 |
0.902 |
8t41-assembly1 |
867 |
143.60 |
0.919 |
8wnj-assembly1 |
920 |
164.12 |
0.897 |
T1125 |
1199 |
196.56 |
0.361 |
T1154 |
1423 |
279.84 |
0.310 |
8uxt-assembly1 |
1596 |
425.98 |
0.783 |
Mean |
— |
112.57 |
0.767 |
Test ID |
Seq Length |
predict_time (s) |
lddt |
|---|---|---|---|
7pv5-assembly1 |
81 |
14.52 |
0.773 |
T1104 |
115 |
19.89 |
0.853 |
T1106s1 |
120 |
20.81 |
0.746 |
7qsj-assembly1 |
373 |
100.22 |
0.936 |
T1137s1 |
408 |
106.16 |
0.797 |
T1112 |
459 |
126.12 |
0.836 |
T1114s3 |
534 |
161.73 |
0.889 |
7uww-assembly1 |
635 |
207.60 |
0.873 |
T1129s2 |
639 |
207.06 |
0.644 |
8k7x-assembly1 |
856 |
384.91 |
0.893 |
8t41-assembly1 |
867 |
406.31 |
0.919 |
8wnj-assembly1 |
920 |
485.19 |
0.897 |
T1125 |
1199 |
948.91 |
0.337 |
T1154 |
1423 |
1782.47 |
0.362 |
8uxt-assembly1 |
1596 |
2430.41 |
0.780 |
Mean |
— |
493.49 |
0.769 |
Version 2.5.0#
This section reports OpenFold2 v2.5.0 performance using internal benchmark artifacts. v2.5.0 adds support for NVIDIA B300 and GB300 GPUs.
Benchmark Configuration#
Parameter |
Setting |
|---|---|
selected_models |
[1,2,3,4,5] |
num_trials |
2 |
The benchmark set is the same four protein chains used in earlier releases:
Test ID |
Seq Length |
|---|---|
7WBN_A |
98 |
7ONG_A |
304 |
7ZHT_A |
562 |
7Y4I_A |
914 |
Table 1: Performance Across the Supported NVIDIA Hardware Units#
The table below reports pipeline_time_mean (seconds) without structural templates. The vs 2.4.0 column is the ratio of v2.4.0 total time to v2.5.0 total time across all four benchmark chains; values above 1.0× indicate faster execution in v2.5.0. B300 and GB300 are new in this release and have no prior baseline.
Hardware |
7WBN_A (98) |
7ONG_A (304) |
7ZHT_A (562) |
7Y4I_A (914) |
vs 2.4.0 |
|---|---|---|---|---|---|
NVIDIA A100 80GB |
5.77 |
27.67 |
61.48 |
131.87 |
1.26× |
NVIDIA B200 |
4.70 |
15.65 |
36.75 |
78.95 |
1.34× |
NVIDIA B300 |
4.78 |
20.26 |
44.05 |
94.02 |
— |
NVIDIA H100 80GB HBM3 |
4.78 |
19.37 |
40.05 |
82.13 |
1.12× |
NVIDIA H200 |
3.99 |
18.26 |
37.75 |
78.23 |
1.04× |
NVIDIA GB200 |
8.53 |
12.30 |
27.81 |
60.60 |
1.36× |
NVIDIA GB300 |
8.73 |
12.07 |
27.15 |
58.94 |
— |
NVIDIA GH200 |
6.96 |
12.93 |
25.97 |
56.44 |
1.16× |
NVIDIA L40S |
5.45 |
26.49 |
64.64 |
145.62 |
1.30× |
NVIDIA GB10 (DGX Spark) |
17.69 |
78.28 |
185.84 |
509.61 |
2.15× |
NVIDIA RTX 6000 Ada |
5.70 |
34.56 |
83.27 |
178.84 |
1.03× |
NVIDIA RTX PRO 6000 Blackwell |
3.90 |
20.35 |
47.78 |
101.42 |
1.47× |
Performance Optimization Tips#
GPU Selection: H100, H200, B200, B300, GB200, and GB300 GPUs deliver the best end-to-end latency for OpenFold2. A100, GH200, RTX PRO 6000 Blackwell, and L40S GPUs are also fully supported.
Sequence Length: The NIM supports sequences from 4 to 2048 residues on a single GPU (1536 on GB10 DGX Spark). Pipeline time scales with sequence length.
Structural Templates: Templates add modest overhead (~1–10s) and can improve prediction accuracy. Refer to Template Processing for guidance.
Memory Management: Ensure adequate GPU memory for your target sequence lengths. Sequences longer than ~1800 residues benefit from 80GB-class or larger GPUs (A100 80GB, H100, H200, B200, B300, GB200, GB300).
Version 2.4.0#
This section reports OpenFold2 v2.4.0 performance using internal benchmark artifacts.
The benchmark set contains four protein chains:
Test ID |
Seq Length |
|---|---|
7WBN_A |
98 |
7ONG_A |
304 |
7ZHT_A |
562 |
7Y4I_A |
914 |
Benchmark Configuration#
Parameter |
Setting |
|---|---|
selected_models |
[1,2,3,4,5] |
num_trials |
2 |
Table 1: Performance Across the Supported NVIDIA Hardware Units#
The table below reports pipeline_time_mean (seconds) with TensorRT backend and no structural templates.
Hardware |
7WBN_A (98) |
7ONG_A (304) |
7ZHT_A (562) |
7Y4I_A (914) |
|---|---|---|---|---|
NVIDIA A100 80GB |
7.08 |
34.16 |
73.82 |
171.75 |
NVIDIA B200 |
5.68 |
22.56 |
48.49 |
104.95 |
NVIDIA H100 80GB HBM3 |
4.60 |
20.03 |
42.66 |
97.04 |
NVIDIA GB200 |
7.15 |
18.16 |
39.93 |
83.19 |
NVIDIA H200 |
4.35 |
16.20 |
37.60 |
85.28 |
NVIDIA L40S |
12.67 |
36.40 |
81.56 |
183.18 |
NVIDIA GB10 (DGX Spark) |
29.70 |
137.10 |
402.87 |
1131.11 |
NVIDIA RTX 6000 Ada |
11.28 |
33.56 |
80.64 |
187.23 |
NVIDIA RTX PRO 6000 Blackwell |
5.68 |
27.59 |
66.81 |
154.09 |
NVIDIA GH200 |
5.63 |
13.81 |
29.65 |
69.50 |
Table 2: Performance Across Optimization Backends#
The table below compares H100 performance between PyTorch and TensorRT backends without structural templates.
Test ID |
Seq Length |
torch (s) |
trt (s) |
trt-speedup-over-torch |
|---|---|---|---|---|
7WBN_A |
98 |
29.74 |
4.60 |
6.47x |
7ONG_A |
304 |
53.34 |
20.03 |
2.66x |
7ZHT_A |
562 |
98.38 |
42.66 |
2.31x |
7Y4I_A |
914 |
199.99 |
97.04 |
2.06x |
Table 3: Performance Impact From Structural Templates#
The table below reports H100 TensorRT performance with and without structural templates.
Test ID |
Seq Length |
Without structural templates (s) |
With structural templates (s) |
|---|---|---|---|
7WBN_A |
98 |
4.60 |
5.29 |
7ONG_A |
304 |
20.03 |
19.53 |
7ZHT_A |
562 |
42.66 |
43.06 |
7Y4I_A |
914 |
97.04 |
97.77 |
Version 2.3.0#
Version 2.3.0 adds support for GB10 (DGX Spark) GPU architecture with optimized performance for this platform.
Performance on GB10 (DGX Spark)#
Below are benchmark times, measured for each input chain, in sequential execution on a single NVIDIA GB10 (DGX Spark) device.
protein chain id |
metric |
7WBN_A |
7ONG_A |
7ZHT_A |
7Y4I_A |
|---|---|---|---|---|---|
sequence length |
sequence length |
98 |
304 |
562 |
914 |
Version 2.3.0 |
pipeline_time |
46.51 |
170.78 |
441.02 |
1079.12 |
Version 2.3.0 |
pipeline_time_per_model |
9.30 |
34.16 |
88.20 |
215.82 |
*pipeline_time is defined as the time to load parameter sets, compute features, and the sum of the time to complete the forward pass for each model in [model_1, model_2, model_3, model_4, model_5]
**pipeline_time divided by 5
Accuracy on GB10#
Accuracy metrics for GB10 (DGX Spark) remain consistent with other supported GPU architectures. For detailed accuracy benchmarks, refer to the Accuracy Metrics in the Version 2.0.0 section below.
For performance benchmarks on other GPU architectures, refer to the Version 2.0.0 section below.
Version 2.2.0#
Version 2.2.0 maintains the same performance characteristics as Version 2.1.0. All performance benchmarks and accuracy metrics from Version 2.1.0 apply to Version 2.2.0.
Version 2.1.0#
Version 2.2.0 maintains the same performance characteristics as version 2.1.0. All performance benchmarks and accuracy metrics from version 2.1.0 apply to version 2.2.0.
For detailed performance comparisons and benchmarks, refer to the Version 2.0.0 section.
Version 2.0.0#
Compared to Version 1.0.0, Version 2.0.0
Removes HHR-based template processing and
Adds TensorRT support for a wide set of GPU architectures. Refer to the Release Notes.
Adds support for ‘explicit_templates’. Refer to the Migrate from HHR-based Templates to Explicit mmCIF Templates.
Performance and Accuracy#
Faster startup: Reduced initialization time due to removal of large template database loading (~300GB)
Enhanced GPU support: TensorRT optimization for L40S, B200, and RTX 6000 Ada Generation
Reduced storage footprint: Container and cache requirements significantly reduced (from 380GB to 80GB total)
Performance Comparison (Time in seconds) on H100#
Below are benchmark times, measured for each input chain, in sequential execution on a single NVIDIA H100 80GB HBM3 device.
protein chain id |
metric |
7WBN_A |
7ONG_A |
7ZHT_A |
7Y4I_A |
|---|---|---|---|---|---|
sequence length |
sequence length |
98 |
304 |
562 |
914 |
Version 1.0.0 |
pipeline_time |
26.6 |
50.9 |
98.2 |
205.9 |
Version 1.0.0 |
pipeline_time_per_model |
5.3 |
10.2 |
19.6 |
41.2 |
Version 2.0.0 |
pipeline_time |
4.92 |
19.4 |
44.9 |
113 |
Version 2.0.0 |
pipeline_time_per_model |
0.984 |
3.90 |
8.99 |
22.6 |
Version 2.0.0 |
speed-up vs 1.0.0 |
5.4x |
2.6x |
2.2x |
1.8x |
*pipeline_time is defined as the time to load parameter sets, compute features, and the sum of the time to complete the forward pass for each model in [model_1, model_2, model_3, model_4, model_5]
**pipeline_time divided by 5
Accuracy Metrics on H100#
Metric |
Version 1.0.0 |
Version 2.0.0 |
|---|---|---|
CADS |
0.744 |
0.740 |
LDDT |
0.861 |
0.860 |
STRIDE |
4.24 |
4.24 |
MP |
0.882 |
0.878 |
Version 1.0.0#
Performance also varies significantly depending on:
The type of NVIDIA GPUs that are attached and available to the NIM
The CPU type
System RAM available
The following section details some performance expectations and provides general tips. These are not meant to be indicative of expected performance and performance on your system varies from these values.
Recommended System Requirements#
Refer to Supported Hardware for requirements.
Performance Benchmarks#
The following are performance benchmarks for OpenFold2 version 1.0.0.
Structure prediction performance is mostly dependent on GPU capability and memory. If you find structure prediction to be a bottleneck, consider using a higher memory device.
The time required for structure prediction grows with sequence length.
The time required for structure prediction grows with the total number of sequences in the alignments.
Below are benchmark times, measured for each input chain, in sequential execution on a single NVIDIA H100 80GB HBM3 device.
The average value of LDDT-CA, for these protein chains, averaged over 2 runs, is 0.86.
protein chain id |
7WBN_A |
7ONG_A |
7ZHT_A |
7Y4I_A |
|---|---|---|---|---|
sequence length |
98 |
304 |
562 |
914 |
pipeline_time* |
26.6 |
50.9 |
98.2 |
205.9 |
pipeline_time_per_model** |
5.3 |
10.2 |
19.6 |
41.2 |
*pipeline_time is defined as the time to load parameter sets, compute features, and the sum of the time to complete the forward pass for each model in [model_1, model_2, model_3, model_4, model_5]
**pipeline_time divided by 5
Version 1.0.0 Configuration#
algo feature / parameter |
setting |
|---|---|
use_templates |
False |
selected_models |
[1,2,3,4,5] |
relax_prediction |
False |
deepspeed evoformer kernel |
active |
precision for deepspeed evoformer kernel |
bf16 |
precision for the rest of the model |
fp32 |