Benchmarks

View as Markdown

BioNeMo Inference Runtime (BioIR) measured against two OSS PyTorch baselines on every GPU it is tested on. These are the numbers a run of the published procedure should reproduce; they are not peak performance, and no attempt was made to tune any configuration per GPU.

Both baselines are reported for every measurement, and the pair is the point. Against OSS PyTorch eager, BioIR is faster on every GPU and model here, by 1.19x at worst. Against OSS torch.compile the answer depends on the architecture: 1.5-3.0x on Ampere, Hopper and Ada, but 1.0-2.1x on Blackwell, where torch.compile closes most of the gap. Either baseline quoted alone overstates the result in one direction or understates it in the other, so figures below 1.00x are printed rather than dropped.

torch.compile is ahead on 19% of individual measurements, and where it is ahead is not random: never below 256 residues, 16% of inputs between 257 and 768, and around 29% above 769 — and only ever on Boltz-2 and OpenFold3, never on OpenFold2 or Protenix. Long inputs on Blackwell are the case to check before quoting a single figure.

Speedup against input size on H100

Per-sample speedup against input size, all four models on H100. The dashed line is parity; crosses mark the largest input measured. The curves are far from flat — high on the smallest inputs, where launch and dispatch overhead dominates and BioIR has less of it, lowest in the few-hundred-residue range, then climbing with length as the fused kernels start to matter. That shape is why every table below is broken out by residue count, and why an overall geomean alone is not enough to quote.

Contents

Measurements

Hardware

LabelDeviceSMMemoryPower limitDriver
h100NVIDIA H100 80GB HBM39080 GiB700 W580.105.08
h200NVIDIA H20090140 GiB700 W595.58.03
a100NVIDIA A100-SXM4-80GB8080 GiB400 W595.58.03
b200NVIDIA B200100179 GiB1000 W595.58.03
b300NVIDIA B300 SXM6 AC103269 GiB1100 W595.58.03
gb200NVIDIA GB200100185 GiB1200 W595.58.03
gb300NVIDIA GB300103278 GiB1400 W595.58.03
l40sNVIDIA L40S8945 GiB350 W595.58.03
rtx6000adaNVIDIA RTX 6000 Ada Generation8948 GiB300 W595.58.03
rtxpro6000svNVIDIA RTX PRO 6000 Blackwell Server Edition12096 GiB600 W595.58.03
rtxpro6000wsNVIDIA RTX PRO 6000 Blackwell Workstation Edition12096 GiB600 W595.58.03
gb10NVIDIA GB10121580.95.05

One GPU per measurement, one process, no replicas. Absolute latencies move with clocks, power limit and thermal headroom; the ratios are what should be comparable.

What Is Compared

Three backends fold the same inputs on the same GPU in the same container:

NameWhat it is
BioIRThis runtime: fused kernels, torch.compile where it applies
OSS torch.compileThe reference implementation under torch.compile
OSS PyTorch eagerThe reference implementation, unmodified

The OSS side runs its own inference script: eager always, and torch.compile additionally when it passes a dynamic-shape probe. Protenix has no compile path, which is why its compile column is empty.

Units

Speedup is OSS forward / BioIR forward per sample; above 1 favours BioIR. Every figure is a geometric mean of per-sample speedups, not a ratio of totals, so each sample counts equally regardless of how long it runs. n is the sample count and differs by model.

Speedups are broken out by input size in residues because the spread is wide — Boltz-2 on H200 is 1.80x below 512 residues and 3.61x above 1024 against OSS torch.compile. The overall geomean alone hides a factor of two.

Peak memory is torch.cuda.max_memory_allocated over the run. Accuracy is OpenStructure lDDT over every sample and DockQ over the subset with a supported protein interface.

Boltz-2

GPUvs OSS torch.compilevs OSS PyTorch eager<512 res512-1024 res>1024 resn
h1002.92x2.66x1.84x / 2.31x3.44x / 2.81x4.17x / 2.94x17
h2001.74x2.54x1.70x / 2.29x1.69x / 2.63x1.87x / 2.76x17
a1003.00x2.80x1.99x / 2.63x3.41x / 2.82x4.23x / 3.01x17
b2001.42x1.46x1.39x / 1.71x1.16x / 1.33x1.84x / 1.36x17
b3001.58x1.62x1.83x / 2.15x1.19x / 1.39x1.89x / 1.38x17
gb2001.13x1.68x1.75x / 2.47x0.90x / 1.36x0.87x / 1.36x17
gb3001.17x1.71x1.77x / 2.56x0.93x / 1.39x0.92x / 1.37x17
l40s1.88x2.78x1.77x / 2.39x1.84x / 2.85x2.04x / 3.23x17
rtx6000ada1.74x2.46x1.81x / 2.27x1.64x / 2.40x1.81x / 2.78x17
rtxpro6000sv1.03x1.50x1.34x / 1.76x0.88x / 1.29x0.90x / 1.48x17
rtxpro6000ws1.06x1.54x1.33x / 1.90x0.93x / 1.30x0.95x / 1.46x17

Bucket cells are OSS torch.compile / OSS PyTorch eager. Below 1.00x means the baseline was faster.

Forward Latency

GPUBioIROSS torch.compileOSS PyTorch eager
h1008.35 s24.36 s22.18 s
h2007.68 s13.39 s19.50 s
a10014.89 s44.70 s41.75 s
b2009.83 s13.95 s14.39 s
b30010.06 s15.92 s16.28 s
gb2009.83 s11.09 s16.49 s
gb3009.59 s11.17 s16.44 s
l40s16.82 s31.54 s46.73 s
rtx6000ada17.13 s29.88 s42.07 s
rtxpro6000sv15.42 s15.84 s23.12 s
rtxpro6000ws14.13 s14.98 s21.74 s

Geomean of the per-sample forward time, over the samples the ratios above are taken over. Quote these when comparing two releases: a ratio that moved does not say which side did.

Peak Allocated Memory

GPUBioIROSS torch.compileOSS PyTorch eagerBioIR / compile
h10020.34 GiB32.27 GiB32.27 GiB0.63x
h20020.34 GiB32.27 GiB32.27 GiB0.63x
a10020.30 GiB32.25 GiB32.25 GiB0.63x
b20019.59 GiB32.27 GiB32.27 GiB0.61x
b30019.59 GiB32.27 GiB32.27 GiB0.61x
gb20019.59 GiB32.27 GiB32.27 GiB0.61x
gb30019.59 GiB32.27 GiB32.27 GiB0.61x
l40s20.30 GiB32.25 GiB32.25 GiB0.63x
rtx6000ada20.30 GiB32.25 GiB32.25 GiB0.63x
rtxpro6000sv19.59 GiB32.27 GiB32.27 GiB0.61x
rtxpro6000ws19.59 GiB32.27 GiB32.27 GiB0.61x

Below 1.00x means BioIR allocates less than the baseline; above 1.00x means it allocates more.

Accuracy

GPUlDDTDockQn lDDTn DockQ
h1000.667 / 0.663 / 0.6670.611 / 0.609 / 0.612179
h2000.667 / 0.664 / 0.6630.611 / 0.609 / 0.617179
a1000.662 / 0.660 / 0.6630.613 / 0.608 / 0.610179
b2000.663 / 0.661 / 0.6630.613 / 0.611 / 0.610179
b3000.669 / 0.660 / 0.6640.606 / 0.601 / 0.605179
gb2000.665 / 0.665 / 0.6620.610 / 0.615 / 0.607179
gb3000.669 / 0.667 / 0.6650.608 / 0.614 / 0.611179
l40s0.664 / 0.667 / 0.6660.603 / 0.612 / 0.615179
rtx6000ada0.663 / 0.665 / 0.6630.613 / 0.616 / 0.608179
rtxpro6000sv0.668 / 0.661 / 0.6660.610 / 0.603 / 0.610179
rtxpro6000ws0.667 / 0.661 / 0.6680.609 / 0.609 / 0.619179

Each cell is BioIR / OSS torch.compile / OSS PyTorch eager. lDDT is scored over every sample; DockQ covers the subset with a supported protein interface, which is why its count is lower. The three should agree — a speedup that moved them would not be the same answer.

OpenFold3

GPUvs OSS torch.compilevs OSS PyTorch eager<512 res512-1024 res>1024 resn
h1001.55x2.02x2.01x / 2.38x1.25x / 1.76x1.48x / 1.96x17
h2001.54x2.03x2.01x / 2.37x1.24x / 1.78x1.46x / 1.97x17
a1001.56x2.02x1.98x / 2.52x1.24x / 1.66x1.54x / 1.95x17
b2001.18x1.47x1.82x / 2.10x0.92x / 1.21x0.94x / 1.22x17
b3001.06x1.40x1.57x / 1.88x0.84x / 1.18x0.86x / 1.21x17
gb2001.38x1.67x2.58x / 2.82x1.08x / 1.29x0.88x / 1.21x17
gb3001.33x1.64x2.46x / 2.73x1.04x / 1.28x0.86x / 1.21x17
l40s1.63x1.90x2.03x / 2.44x1.39x / 1.62x1.51x / 1.70x17
rtx6000ada1.63x1.96x1.95x / 2.37x1.41x / 1.72x1.54x / 1.82x17
rtxpro6000sv1.07x1.29x1.64x / 1.92x0.83x / 1.03x0.86x / 1.04x17
rtxpro6000ws1.06x1.31x1.55x / 1.98x0.84x / 1.03x0.90x / 1.05x17

Bucket cells are OSS torch.compile / OSS PyTorch eager. Below 1.00x means the baseline was faster.

Forward Latency

GPUBioIROSS torch.compileOSS PyTorch eager
h10011.12 s17.26 s22.46 s
h20010.23 s15.78 s20.76 s
a10018.93 s29.49 s38.21 s
b20011.87 s14.00 s17.49 s
b30012.14 s12.83 s17.00 s
gb20011.59 s16.03 s19.34 s
gb30011.64 s15.51 s19.13 s
l40s21.55 s35.13 s40.91 s
rtx6000ada20.75 s33.73 s40.68 s
rtxpro6000sv17.95 s19.16 s23.07 s
rtxpro6000ws16.48 s17.53 s21.56 s

Geomean of the per-sample forward time, over the samples the ratios above are taken over. Quote these when comparing two releases: a ratio that moved does not say which side did.

Peak Allocated Memory

GPUBioIROSS torch.compileOSS PyTorch eagerBioIR / compile
h10022.14 GiB29.49 GiB29.49 GiB0.75x
h20022.14 GiB29.49 GiB29.49 GiB0.75x
a10022.10 GiB29.47 GiB29.47 GiB0.75x
b20021.39 GiB29.49 GiB29.49 GiB0.73x
b30021.39 GiB29.49 GiB29.49 GiB0.73x
gb20021.39 GiB29.49 GiB29.49 GiB0.73x
gb30021.39 GiB29.49 GiB29.49 GiB0.73x
l40s22.10 GiB29.47 GiB29.47 GiB0.75x
rtx6000ada22.10 GiB29.47 GiB29.47 GiB0.75x
rtxpro6000sv21.39 GiB29.49 GiB29.49 GiB0.73x
rtxpro6000ws21.39 GiB29.49 GiB29.49 GiB0.73x

Below 1.00x means BioIR allocates less than the baseline; above 1.00x means it allocates more.

Accuracy

GPUlDDTDockQn lDDTn DockQ
h1000.617 / 0.616 / 0.6190.533 / — / —179
h2000.617 / 0.617 / 0.6120.533 / — / —179
a1000.616 / 0.616 / 0.6180.533 / — / —179
b2000.618 / 0.616 / 0.6140.535 / — / —179
b3000.618 / 0.619 / 0.6170.532 / — / —179
gb2000.613 / 0.606 / 0.6130.531 / — / —179
gb3000.614 / 0.619 / 0.6180.530 / — / —179
l40s0.617 / 0.617 / 0.6150.544 / — / —179
rtx6000ada0.614 / 0.621 / 0.6070.533 / — / —179
rtxpro6000sv0.618 / 0.617 / 0.6180.538 / — / —179
rtxpro6000ws0.617 / 0.620 / 0.6180.533 / — / —179

Each cell is BioIR / OSS torch.compile / OSS PyTorch eager. lDDT is scored over every sample; DockQ covers the subset with a supported protein interface, which is why its count is lower. The three should agree — a speedup that moved them would not be the same answer.

OpenFold2 / AlphaFold2

Monomer

GPUvs OSS torch.compilevs OSS PyTorch eager<512 res512-1024 res>1024 resn
h1002.55x2.60x2.17x / 2.11x2.67x / 2.83x3.23x / 3.35x15
h2002.61x2.66x2.24x / 2.16x2.73x / 2.90x3.25x / 3.37x15
a1002.58x2.71x2.15x / 2.30x2.70x / 2.84x3.35x / 3.46x15
b2001.82x1.90x1.73x / 1.72x1.84x / 2.00x1.99x / 2.07x15
b3001.83x1.88x1.74x / 1.69x1.84x / 1.99x2.04x / 2.05x15
gb2001.94x2.03x1.77x / 1.79x2.02x / 2.19x2.16x / 2.26x15
gb3002.05x2.03x1.91x / 1.79x2.13x / 2.18x2.21x / 2.25x15
l40s1.78x1.77x1.66x / 1.57x1.72x / 1.78x2.22x / 2.27x15
rtx6000ada2.01x2.08x1.90x / 1.94x1.98x / 2.08x2.30x / 2.36x15
rtxpro6000sv1.72x1.71x1.73x / 1.62x1.63x / 1.69x1.88x / 1.92x15
rtxpro6000ws1.72x1.77x1.72x / 1.75x1.65x / 1.72x1.87x / 1.91x15
gb101.31x1.36x1.16x / 1.28x1.60x / 1.60x1.10x / 1.12x15

Bucket cells are OSS torch.compile / OSS PyTorch eager. Below 1.00x means the baseline was faster.

Forward Latency

GPUBioIROSS torch.compileOSS PyTorch eager
h1003.05 s7.78 s7.93 s
h2002.89 s7.56 s7.69 s
a1005.33 s13.74 s14.46 s
b2003.51 s6.40 s6.67 s
b3003.38 s6.20 s6.34 s
gb2003.66 s7.12 s7.44 s
gb3003.70 s7.59 s7.49 s
l40s6.59 s11.74 s11.69 s
rtx6000ada6.43 s12.91 s13.34 s
rtxpro6000sv4.42 s7.60 s7.54 s
rtxpro6000ws4.12 s7.08 s7.29 s
gb1029.04 s37.91 s39.57 s

Geomean of the per-sample forward time, over the samples the ratios above are taken over. Quote these when comparing two releases: a ratio that moved does not say which side did.

Peak Allocated Memory

GPUBioIROSS torch.compileOSS PyTorch eagerBioIR / compile
h10029.69 GiB15.27 GiB15.27 GiB1.94x
h20029.69 GiB15.27 GiB15.27 GiB1.94x
a10029.67 GiB15.25 GiB15.25 GiB1.95x
b20018.77 GiB15.27 GiB15.27 GiB1.23x
b30018.77 GiB15.27 GiB15.27 GiB1.23x
gb20018.77 GiB15.27 GiB15.27 GiB1.23x
gb30018.77 GiB15.27 GiB15.27 GiB1.23x
l40s29.67 GiB15.25 GiB15.25 GiB1.95x
rtx6000ada29.67 GiB15.25 GiB15.25 GiB1.95x
rtxpro6000sv23.56 GiB15.27 GiB15.27 GiB1.54x
rtxpro6000ws23.56 GiB15.27 GiB15.27 GiB1.54x
gb1023.56 GiB15.27 GiB15.27 GiB1.54x

Below 1.00x means BioIR allocates less than the baseline; above 1.00x means it allocates more.

Accuracy

GPUlDDTDockQn lDDTn DockQ
h1000.434 / 0.426 / 0.432— / — / —15
h2000.434 / 0.421 / 0.426— / — / —15
a1000.434 / 0.422 / 0.422— / — / —15
b2000.434 / 0.429 / 0.433— / — / —15
b3000.434 / 0.427 / 0.426— / — / —15
gb2000.433 / 0.432 / 0.429— / — / —15
gb3000.433 / 0.428 / 0.423— / — / —15
l40s0.432 / 0.431 / 0.426— / — / —15
rtx6000ada0.434 / 0.420 / 0.418— / — / —15
rtxpro6000sv0.435 / 0.429 / 0.418— / — / —15
rtxpro6000ws0.435 / 0.426 / 0.433— / — / —15
gb100.434 / 0.423 / 0.424— / — / —15

Each cell is BioIR / OSS torch.compile / OSS PyTorch eager. lDDT is scored over every sample; DockQ covers the subset with a supported protein interface, which is why its count is lower. The three should agree — a speedup that moved them would not be the same answer.

Multimer

GPUvs OSS torch.compilevs OSS PyTorch eager<512 res512-1024 res>1024 resn
h1002.66x2.77x— / —2.40x / 2.54x2.94x / 3.03x4
h2002.61x2.75x— / —2.35x / 2.51x2.90x / 3.00x4
a1002.64x2.74x— / —2.41x / 2.53x2.89x / 2.96x4
b2001.80x1.90x— / —1.68x / 1.82x1.92x / 2.00x4
b3001.79x1.89x— / —1.68x / 1.80x1.92x / 1.99x4
gb2001.80x1.91x— / —1.68x / 1.81x1.93x / 2.01x4
gb3001.80x1.90x— / —1.68x / 1.81x1.92x / 1.99x4
l40s2.13x2.18x— / —1.86x / 1.92x2.44x / 2.47x4
rtx6000ada1.96x2.03x— / —1.78x / 1.87x2.16x / 2.21x4
rtxpro6000sv1.83x1.87x— / —1.59x / 1.64x2.10x / 2.14x4
rtxpro6000ws1.76x1.81x— / —1.55x / 1.62x2.00x / 2.03x4
gb101.41x1.45x— / —1.68x / 1.72x1.19x / 1.22x4

Bucket cells are OSS torch.compile / OSS PyTorch eager. Below 1.00x means the baseline was faster.

Forward Latency

GPUBioIROSS torch.compileOSS PyTorch eager
h1008.15 s21.64 s22.59 s
h2007.69 s20.10 s21.12 s
a10014.66 s38.68 s40.11 s
b2009.35 s16.78 s17.80 s
b3009.08 s16.29 s17.16 s
gb2008.73 s15.74 s16.66 s
gb3008.69 s15.62 s16.50 s
l40s19.23 s40.99 s41.91 s
rtx6000ada21.33 s41.90 s43.35 s
rtxpro6000sv12.18 s22.29 s22.83 s
rtxpro6000ws12.34 s21.71 s22.33 s
gb10101.50 s143.55 s146.83 s

Geomean of the per-sample forward time, over the samples the ratios above are taken over. Quote these when comparing two releases: a ratio that moved does not say which side did.

Peak Allocated Memory

GPUBioIROSS torch.compileOSS PyTorch eagerBioIR / compile
h10018.01 GiB14.04 GiB14.04 GiB1.28x
h20018.01 GiB14.04 GiB14.04 GiB1.28x
a10017.99 GiB14.02 GiB14.02 GiB1.28x
b20016.31 GiB14.04 GiB14.04 GiB1.16x
b30016.31 GiB14.04 GiB14.04 GiB1.16x
gb20016.31 GiB14.04 GiB14.04 GiB1.16x
gb30016.31 GiB14.04 GiB14.04 GiB1.16x
l40s17.99 GiB14.02 GiB14.02 GiB1.28x
rtx6000ada17.99 GiB14.02 GiB14.02 GiB1.28x
rtxpro6000sv21.43 GiB14.04 GiB14.04 GiB1.53x
rtxpro6000ws21.43 GiB14.04 GiB14.04 GiB1.53x
gb1021.43 GiB14.04 GiB14.04 GiB1.53x

Below 1.00x means BioIR allocates less than the baseline; above 1.00x means it allocates more.

Accuracy

GPUlDDTDockQn lDDTn DockQ
h1000.784 / 0.736 / 0.7660.603 / 0.601 / 0.61044
h2000.784 / 0.762 / 0.7610.603 / 0.591 / 0.60644
a1000.777 / 0.750 / 0.7770.597 / 0.617 / 0.62044
b2000.783 / 0.766 / 0.7740.607 / 0.610 / 0.60644
b3000.783 / 0.756 / 0.7790.607 / 0.609 / 0.62044
gb2000.783 / 0.768 / 0.7620.607 / 0.600 / 0.61044
gb3000.783 / 0.768 / 0.7740.607 / 0.635 / 0.62144
l40s0.780 / 0.783 / 0.7660.609 / 0.613 / 0.60044
rtx6000ada0.782 / 0.769 / 0.7660.602 / 0.622 / 0.60844
rtxpro6000sv0.776 / 0.746 / 0.7620.601 / 0.594 / 0.60844
rtxpro6000ws0.776 / 0.758 / 0.7590.601 / 0.595 / 0.60344
gb100.780 / 0.765 / 0.7340.602 / 0.600 / 0.60344

Each cell is BioIR / OSS torch.compile / OSS PyTorch eager. lDDT is scored over every sample; DockQ covers the subset with a supported protein interface, which is why its count is lower. The three should agree — a speedup that moved them would not be the same answer.

Protenix

GPUvs OSS torch.compilevs OSS PyTorch eager<512 res512-1024 res>1024 resn
h1001.87x— / 2.06x— / 1.66x— / 1.92x17
h2001.84x— / 2.03x— / 1.66x— / 1.83x17
a1001.88x— / 2.21x— / 1.62x— / 1.85x17
b2001.27x— / 1.85x— / 1.10x— / 0.96x17
b3001.21x— / 1.59x— / 1.12x— / 0.96x17
gb2001.49x— / 2.62x— / 1.20x— / 0.97x17
gb3001.48x— / 2.57x— / 1.20x— / 0.98x17
l40s1.76x— / 1.98x— / 1.64x— / 1.66x17
rtx6000ada1.68x— / 2.26x— / 1.42x— / 1.45x17
rtxpro6000sv1.19x— / 1.61x— / 1.01x— / 1.00x17
rtxpro6000ws1.21x— / 1.68x— / 1.00x— / 1.01x17

Bucket cells are OSS torch.compile / OSS PyTorch eager. Below 1.00x means the baseline was faster.

Forward Latency

GPUBioIROSS torch.compileOSS PyTorch eager
h10013.05 s24.39 s
h20012.18 s22.37 s
a10023.09 s43.40 s
b20013.98 s17.75 s
b30014.26 s17.27 s
gb20013.73 s20.43 s
gb30013.61 s20.12 s
l40s25.89 s45.56 s
rtx6000ada30.37 s51.15 s
rtxpro6000sv21.13 s25.14 s
rtxpro6000ws19.67 s23.71 s

Geomean of the per-sample forward time, over the samples the ratios above are taken over. Quote these when comparing two releases: a ratio that moved does not say which side did.

Peak Allocated Memory

GPUBioIROSS torch.compileOSS PyTorch eagerBioIR / compile
h10027.57 GiB35.16 GiB
h20027.57 GiB35.16 GiB
a10027.55 GiB35.14 GiB
b20033.19 GiB35.16 GiB
b30033.19 GiB35.16 GiB
gb20033.19 GiB35.16 GiB
gb30033.19 GiB35.16 GiB
l40s27.55 GiB35.14 GiB
rtx6000ada27.55 GiB35.14 GiB
rtxpro6000sv33.19 GiB35.16 GiB
rtxpro6000ws33.19 GiB35.16 GiB

Below 1.00x means BioIR allocates less than the baseline; above 1.00x means it allocates more.

Accuracy

GPUlDDTDockQn lDDTn DockQ
h1000.658 / — / 0.6500.591 / — / 0.545179
h2000.657 / — / 0.6500.591 / — / 0.568179
a1000.659 / — / 0.6490.599 / — / 0.541179
b2000.659 / — / 0.6480.590 / — / 0.562179
b3000.658 / — / 0.6540.589 / — / 0.549179
gb2000.659 / — / 0.6500.588 / — / 0.550179
gb3000.659 / — / 0.6530.589 / — / 0.555179
l40s0.660 / — / 0.6520.600 / — / 0.552179
rtx6000ada0.658 / — / 0.6500.601 / — / 0.554179
rtxpro6000sv0.658 / — / 0.6490.598 / — / 0.548179
rtxpro6000ws0.658 / — / 0.6460.588 / — / 0.542179

Each cell is BioIR / OSS torch.compile / OSS PyTorch eager. lDDT is scored over every sample; DockQ covers the subset with a supported protein interface, which is why its count is lower. The three should agree — a speedup that moved them would not be the same answer.

Reproducing

These numbers are not tied to a build artifact: they come from main, on one GPU, with public inputs and upstream baselines at pinned tags. The procedure is the bench-perf-oss agent skill — point an agent at it rather than following steps by hand. It installs BioIR and each pinned baseline in isolation, runs one warmup and one timed forward per sample, scores the output, and aggregates the geomeans above.

Have these ready first; the skill cannot supply them:

  • One NVIDIA GPU on driver 580 or newer. The container ships a CUDA 13.2 build of PyTorch. On a 535 driver it runs only through the forward-compatibility shim, and measuring on that path produced silent hangs and SIGFPE crashes part-way through a sweep — wrong numbers, not merely slow ones. Check the ECC counters are zero while you are there.
  • NGC_API_KEY exported, for the MSA Search NIM. rebuild_dataset.py builds the dataset-0.3.0 inputs from RCSB and that endpoint rather than downloading a release; the structures need no credential at all (--structures-only). Free key at build.nvidia.com.
  • HF_TOKEN exported. OpenFold3’s checkpoint is gated and an anonymous fetch returns 401. Registering is free. Checkpoint staging generally is Model Weights.
  • A built extension, per Development Workflow.

When a ratio does not match, suspect the baseline environment before the code. The most common cause by far is an upstream package resolved against a different torch than BioIR’s, at which point the ratio compares two PyTorch builds rather than two implementations.


Container nvcr.io/nvidia/pytorch:26.05-py3 · driver 580.105.08–595.58.03 · measured 2026-09-01