Release Notes for NVIDIA NIM for OpenFold3#

Release 1.6.0#

Summary#

This release substantially reduces prediction latency, particularly for short sequences, and adds the ability to serve requests from several GPUs in one container.

Key Features#

  • Faster predictions through CUDA graph capture: The BioNeMo Inference Runtime now captures the diffusion module as a CUDA graph. What this removes is per-request kernel launch overhead, which is a fixed cost rather than a proportional one: across the benchmark suite on H100 it takes roughly 7-8 seconds off every prediction, whatever the input size.

    The speedup therefore looks very different at different sizes, from the same change. A 186-residue prediction drops from 10.2 to 2.9 seconds, because about 7 of those 10 seconds were launch overhead. A 1869-residue prediction drops from 82.5 to 72.5 seconds, the same 10 seconds but now a small fraction of the total. Expect a large improvement on short inputs, a modest one on long inputs, and on the longest inputs on some SKUs no measurable change. Refer to Performance for measured per-SKU results.

  • Multi-GPU serving: A container given more than one GPU starts one inference process per GPU and serves requests concurrently. Each prediction still runs on a single GPU, so this raises throughput rather than lowering single-request latency. Refer to Configuration Alternatives.

  • Bounded request queue: When every GPU is busy and the queue is full, the NIM returns HTTP 503 rather than accepting work it cannot start. Clients should retry after a short delay.

  • Serving your own checkpoint is documented and enforced: MODEL_PATH and MODEL_FILE_NAME select a checkpoint of the same architecture, fine-tuned weights for example. A path that does not resolve is now fatal at startup instead of falling back to the shipped checkpoint, which previously made a misconfigured mount look like a working one. Refer to Configuration Alternatives.

  • Reproducible predictions: The sampler is seeded per request, so the same input returns the same structures on every call. Asking for several diffusion samples still returns several different structures within that response. It is repeat calls that now agree. NIM_INFERENCE_SEED= (empty) restores the previous behaviour, where repeat calls varied.

  • Abandoned requests are discarded: A queued request whose client has disconnected is dropped rather than run, and reported as HTTP 499. Previously one abandoned long prediction could hold the GPU while every request behind it waited for a result nobody would read.

  • Out-of-memory requests are reported, not hidden: A prediction too large for the GPU now returns HTTP 413 naming the residue count, the sample count and how much memory was wanted. It previously returned HTTP 500 with “Internal server error.”, which was indistinguishable from a genuine fault.

  • Updated base image: The container now builds on CUDA 13.2 and PyTorch 2.12. This also reduced the container’s CVE exposure relative to 1.5.0.

  • Smaller container: Roughly 11GB to download, down from the previous release.

Notes and Limitations#

  • Requests are not batched. Concurrency comes from serving one request per GPU; a single request never uses more than one GPU.

  • All features from release 1.5.0 remain supported.

Release 1.5.0#

Summary#

This release adds support for NVIDIA B300 and GB300 GPUs and introduces a new, fully NVIDIA-optimized inference backend that replaces the previous hybrid implementation.

Key Features#

  • NVIDIA B300: Added support for NVIDIA B300 GPU

  • NVIDIA GB300: Added support for NVIDIA GB300 superchip

  • New optimized inference backend: The NIM now runs on a fully NVIDIA-tuned inference path, delivering faster and more consistent predictions out of the box.

Backend Optimizations#

The new inference backend introduces custom GPU kernels and graph-level optimizations:

  • New kernels: Fused Adaptive Layer Norm; DualGemm (sm80, two variants); gated sigmoid (sm80); FlashAttention v2/v3 for Triangle Attention and Attention Pair Bias.

  • Pre-computed masks and buffers: Pre-computed, left-aligned attention masks feed directly into the new dual-GEMM and attention kernels, with pre-computed buffers for the Pairformer and Diffusion Transformer (DiT) stacks.

  • Lower peak memory: The confidence module is evaluated in a loop to reduce memory usage.

  • Diffusion Transformer: Removed the bias computation for Attention Pair Bias, replacing it with three kernels (one reduce, one GEMM, one padding).

  • Optimized query-to-keys path in the attention computation.

Notes and Limitations#

  • The optional PyTorch-only inference mode from earlier releases is no longer available. All requests run on the new optimized backend.

  • All features from release 1.4.0 remain supported.

Release 1.4.0#

Summary#

This release extends GPU support with the addition of NVIDIA GH200 superchips and the RTX PRO 6000 Blackwell Workstation Edition, adds Slurm support for HPC deployments, and introduces a new checkpoint from OpenFold lab that improves prediction accuracy.

Key Features#

  • NVIDIA GH200 144GB: Added support for GH200 with 144GB memory

  • NVIDIA RTX PRO 6000 Blackwell Workstation Edition: Added support for the professional workstation GPU with 96GB GDDR7 memory

  • Slurm support: Added support for running on Slurm-based HPC clusters

  • New checkpoint from OpenFold lab: Added support for a new checkpoint that improves prediction accuracy

Release 1.3.0#

Summary#

This release extends GPU support with the addition of the GB10 architecture and updates cuEquivariance for enhanced performance.

Key Features#

  • GB10 DGX Spark Support: Added support for GB10 DGX Spark SKUs with sequence lengths up to 1536 residues

  • cuEquivariance integration: Updated to cuEquivariance 0.8.1 for GB10 support

Release 1.2.0#

Summary#

This release adds support for GB200 GPU with ARM architecture and extends hardware compatibility for enterprise deployments.

Key Features#

  • GB200 Support: Added support for GB200 GPU

  • Enhanced compatibility with ARM-based systems

Notes and Limitations#

  • Ensure you use this NIM with GPUs with at least 48 GB of VRAM.

  • CUDA Driver Requirements:

    • Minimum version 580 with CUDA 13.0

Note: While there are many options for tuning this NIM’s performance, for most users, the defaults provide a balanced performance experience.

Release 1.1.0#

Summary#

This release adds support for protein structural template inputs for complex structure prediction, and additional GPU hardware.

New Features#

Structural Template Support#

The NIM now accepts protein structural template inputs (single-chain) to guide structure prediction:

  • Template input: For each input, protein molecules can include structural templates in CIF format

  • Template guidance: Templates help constrain and improve structure predictions when experimental or predicted structures are available

  • Multiple templates: Each protein can have multiple structural templates

Additional GPU Support#

Support for additional NVIDIA GPU configurations:

  • NVIDIA B200: High-performance GPU for large-scale structure prediction

  • NVIDIA L40S: Cost-effective GPU option for production deployments

Telemetry Control#

Telemetry control: NIM Telemetry helps NVIDIA deliver a faster, more reliable experience with greater compatibility across a wide range of environments, while maintaining strict privacy protections and giving users full control.

Benefits:

  • Enhances performance and reliability: Provides anonymous system and NIM-level insights that help NVIDIA identify bottlenecks, tune performance across hardware configurations, and improve runtime stability.

  • Improves compatibility across deployments: Helps detect and resolve version, driver, and environment compatibility issues early, reducing friction across diverse infrastructure setups.

  • Accelerates troubleshooting and bug resolution: Allows NVIDIA to diagnose errors and regressions faster, leading to quicker support response times and higher overall availability.

  • Informs smarter optimizations and future releases: Real-world, aggregated telemetry data helps guide the optimization of NIM runtimes, model packaging, and deployment workflows, ensuring updates target the scenarios that matter most to users.

  • Protects user privacy and data security: Collects only minimal, anonymous metadata, such as hardware type and NIM version. No user data, input sequences, or prediction results are collected.

  • Fully optional and configurable: Telemetry collection is disabled by default. You can toggle telemetry at any time using environment variables.

Configuration:

  • Set NIM_TELEMETRY_MODE=0 to disable telemetry (default)

  • Set NIM_TELEMETRY_MODE=1 to enable telemetry

For more information about data privacy, what is collected, and how to configure telemetry, refer to:

Supported Features#

All features from release 1.0.0 remain supported, with the addition of structural templates for protein molecules.

Release 1.0.0#

Summary#

This is the first release of NVIDIA NIM for OpenFold3.

Supported Features#

Entity Types#

The NIM request API accepts the following molecular entity types:

  • Protein: Amino acid sequences

  • DNA: DNA sequences

  • RNA: RNA sequences

  • Ligand: Small molecules specified via CCD codes or SMILES strings

Multiple Entities and Complexes#

The API accepts inputs with multiple entities across multiple types, enabling prediction of complex biomolecular structures. For example:

  • Monomeric proteins: 1 protein sequence

  • Protein-ligand complexes: 1 or more protein sequences, 1 or more ligands

  • Homomeric protein complexes: 2 or more identical protein sequences

  • Heteromeric protein complexes: 2 or more nonidentical protein sequences

  • Protein-DNA complexes: 1 or more protein sequences, 1 or more DNA sequences

  • Protein-RNA complexes: 1 or more protein sequences, 1 or more RNA sequences

Multiple Sequence Alignment (MSA) Support#

The NIM accepts MSA inputs with flexible configuration:

  • MSA Requirements: MSA inputs are required for protein and RNA entity types

  • Format support: Multiple alignment file formats are supported:

    • a3m : A3M format

    • csv : CSV format

  • MSA types:

    • Unpaired MSAs for protein sequences and RNA sequences.

    • Paired MSAs for inputs with multiple non-identical protein sequences.

Note: The OpenFold3 NIM accepts paired MSA inputs, but does not perform ‘online’ pairing of MSAs.

Inference Parameters#

Forward Pass Parameters:

  • diffusion_samples: Number of independent structures to generate (1-5, default: 1)

Structure Templates#

  • Note: This release does not support structure template inputs. Predictions are generated without template guidance.

Output Formats#

  • CIF: Crystallographic Information File format (supported)

  • PDB: Protein Data Bank format (supported)

  • mmCIF: Macromolecular CIF format (not yet supported)