Release Notes for NVIDIA NIM for OpenFold3#
Release 1.6.0#
Summary#
This release substantially reduces prediction latency, particularly for short sequences, and adds the ability to serve requests from several GPUs in one container.
Key Features#
Faster predictions through CUDA graph capture: The BioNeMo Inference Runtime now captures the diffusion module as a CUDA graph. What this removes is per-request kernel launch overhead, which is a fixed cost rather than a proportional one: across the benchmark suite on H100 it takes roughly 7-8 seconds off every prediction, whatever the input size.
The speedup therefore looks very different at different sizes, from the same change. A 186-residue prediction drops from 10.2 to 2.9 seconds, because about 7 of those 10 seconds were launch overhead. A 1869-residue prediction drops from 82.5 to 72.5 seconds, the same 10 seconds but now a small fraction of the total. Expect a large improvement on short inputs, a modest one on long inputs, and on the longest inputs on some SKUs no measurable change. Refer to Performance for measured per-SKU results.
Multi-GPU serving: A container given more than one GPU starts one inference process per GPU and serves requests concurrently. Each prediction still runs on a single GPU, so this raises throughput rather than lowering single-request latency. Refer to Configuration Alternatives.
Bounded request queue: When every GPU is busy and the queue is full, the NIM returns HTTP 503 rather than accepting work it cannot start. Clients should retry after a short delay.
Serving your own checkpoint is documented and enforced:
MODEL_PATHandMODEL_FILE_NAMEselect a checkpoint of the same architecture, fine-tuned weights for example. A path that does not resolve is now fatal at startup instead of falling back to the shipped checkpoint, which previously made a misconfigured mount look like a working one. Refer to Configuration Alternatives.Reproducible predictions: The sampler is seeded per request, so the same input returns the same structures on every call. Asking for several diffusion samples still returns several different structures within that response. It is repeat calls that now agree.
NIM_INFERENCE_SEED=(empty) restores the previous behaviour, where repeat calls varied.Abandoned requests are discarded: A queued request whose client has disconnected is dropped rather than run, and reported as HTTP 499. Previously one abandoned long prediction could hold the GPU while every request behind it waited for a result nobody would read.
Out-of-memory requests are reported, not hidden: A prediction too large for the GPU now returns HTTP 413 naming the residue count, the sample count and how much memory was wanted. It previously returned HTTP 500 with “Internal server error.”, which was indistinguishable from a genuine fault.
Updated base image: The container now builds on CUDA 13.2 and PyTorch 2.12. This also reduced the container’s CVE exposure relative to 1.5.0.
Smaller container: Roughly 11GB to download, down from the previous release.
Notes and Limitations#
Requests are not batched. Concurrency comes from serving one request per GPU; a single request never uses more than one GPU.
All features from release 1.5.0 remain supported.
Release 1.5.0#
Summary#
This release adds support for NVIDIA B300 and GB300 GPUs and introduces a new, fully NVIDIA-optimized inference backend that replaces the previous hybrid implementation.
Key Features#
NVIDIA B300: Added support for NVIDIA B300 GPU
NVIDIA GB300: Added support for NVIDIA GB300 superchip
New optimized inference backend: The NIM now runs on a fully NVIDIA-tuned inference path, delivering faster and more consistent predictions out of the box.
Backend Optimizations#
The new inference backend introduces custom GPU kernels and graph-level optimizations:
New kernels: Fused Adaptive Layer Norm; DualGemm (sm80, two variants); gated sigmoid (sm80); FlashAttention v2/v3 for Triangle Attention and Attention Pair Bias.
Pre-computed masks and buffers: Pre-computed, left-aligned attention masks feed directly into the new dual-GEMM and attention kernels, with pre-computed buffers for the Pairformer and Diffusion Transformer (DiT) stacks.
Lower peak memory: The confidence module is evaluated in a loop to reduce memory usage.
Diffusion Transformer: Removed the bias computation for Attention Pair Bias, replacing it with three kernels (one reduce, one GEMM, one padding).
Optimized query-to-keys path in the attention computation.
Notes and Limitations#
The optional PyTorch-only inference mode from earlier releases is no longer available. All requests run on the new optimized backend.
All features from release 1.4.0 remain supported.
Release 1.4.0#
Summary#
This release extends GPU support with the addition of NVIDIA GH200 superchips and the RTX PRO 6000 Blackwell Workstation Edition, adds Slurm support for HPC deployments, and introduces a new checkpoint from OpenFold lab that improves prediction accuracy.
Key Features#
NVIDIA GH200 144GB: Added support for GH200 with 144GB memory
NVIDIA RTX PRO 6000 Blackwell Workstation Edition: Added support for the professional workstation GPU with 96GB GDDR7 memory
Slurm support: Added support for running on Slurm-based HPC clusters
New checkpoint from OpenFold lab: Added support for a new checkpoint that improves prediction accuracy
Release 1.3.0#
Summary#
This release extends GPU support with the addition of the GB10 architecture and updates cuEquivariance for enhanced performance.
Key Features#
GB10 DGX Spark Support: Added support for GB10 DGX Spark SKUs with sequence lengths up to 1536 residues
cuEquivariance integration: Updated to cuEquivariance 0.8.1 for GB10 support
Release 1.2.0#
Summary#
This release adds support for GB200 GPU with ARM architecture and extends hardware compatibility for enterprise deployments.
Key Features#
GB200 Support: Added support for GB200 GPU
Enhanced compatibility with ARM-based systems
Notes and Limitations#
Ensure you use this NIM with GPUs with at least 48 GB of VRAM.
CUDA Driver Requirements:
Minimum version 580 with CUDA 13.0
Note: While there are many options for tuning this NIM’s performance, for most users, the defaults provide a balanced performance experience.
Release 1.1.0#
Summary#
This release adds support for protein structural template inputs for complex structure prediction, and additional GPU hardware.
New Features#
Structural Template Support#
The NIM now accepts protein structural template inputs (single-chain) to guide structure prediction:
Template input: For each input, protein molecules can include structural templates in CIF format
Template guidance: Templates help constrain and improve structure predictions when experimental or predicted structures are available
Multiple templates: Each protein can have multiple structural templates
Additional GPU Support#
Support for additional NVIDIA GPU configurations:
NVIDIA B200: High-performance GPU for large-scale structure prediction
NVIDIA L40S: Cost-effective GPU option for production deployments
Telemetry Control#
Telemetry control: NIM Telemetry helps NVIDIA deliver a faster, more reliable experience with greater compatibility across a wide range of environments, while maintaining strict privacy protections and giving users full control.
Benefits:
Enhances performance and reliability: Provides anonymous system and NIM-level insights that help NVIDIA identify bottlenecks, tune performance across hardware configurations, and improve runtime stability.
Improves compatibility across deployments: Helps detect and resolve version, driver, and environment compatibility issues early, reducing friction across diverse infrastructure setups.
Accelerates troubleshooting and bug resolution: Allows NVIDIA to diagnose errors and regressions faster, leading to quicker support response times and higher overall availability.
Informs smarter optimizations and future releases: Real-world, aggregated telemetry data helps guide the optimization of NIM runtimes, model packaging, and deployment workflows, ensuring updates target the scenarios that matter most to users.
Protects user privacy and data security: Collects only minimal, anonymous metadata, such as hardware type and NIM version. No user data, input sequences, or prediction results are collected.
Fully optional and configurable: Telemetry collection is disabled by default. You can toggle telemetry at any time using environment variables.
Configuration:
Set
NIM_TELEMETRY_MODE=0to disable telemetry (default)Set
NIM_TELEMETRY_MODE=1to enable telemetry
For more information about data privacy, what is collected, and how to configure telemetry, refer to:
Supported Features#
All features from release 1.0.0 remain supported, with the addition of structural templates for protein molecules.
Release 1.0.0#
Summary#
This is the first release of NVIDIA NIM for OpenFold3.
Supported Features#
Entity Types#
The NIM request API accepts the following molecular entity types:
Protein: Amino acid sequences
DNA: DNA sequences
RNA: RNA sequences
Ligand: Small molecules specified via CCD codes or SMILES strings
Multiple Entities and Complexes#
The API accepts inputs with multiple entities across multiple types, enabling prediction of complex biomolecular structures. For example:
Monomeric proteins: 1 protein sequence
Protein-ligand complexes: 1 or more protein sequences, 1 or more ligands
Homomeric protein complexes: 2 or more identical protein sequences
Heteromeric protein complexes: 2 or more nonidentical protein sequences
Protein-DNA complexes: 1 or more protein sequences, 1 or more DNA sequences
Protein-RNA complexes: 1 or more protein sequences, 1 or more RNA sequences
Multiple Sequence Alignment (MSA) Support#
The NIM accepts MSA inputs with flexible configuration:
MSA Requirements: MSA inputs are required for protein and RNA entity types
Format support: Multiple alignment file formats are supported:
a3m: A3M formatcsv: CSV format
MSA types:
Unpaired MSAs for protein sequences and RNA sequences.
Paired MSAs for inputs with multiple non-identical protein sequences.
Note: The OpenFold3 NIM accepts paired MSA inputs, but does not perform ‘online’ pairing of MSAs.
Inference Parameters#
Forward Pass Parameters:
diffusion_samples: Number of independent structures to generate (1-5, default: 1)
Structure Templates#
Note: This release does not support structure template inputs. Predictions are generated without template guidance.
Output Formats#
CIF: Crystallographic Information File format (supported)
PDB: Protein Data Bank format (supported)
mmCIF: Macromolecular CIF format (not yet supported)