Release Notes#

Release 1.10.0#

Summary#

This release adds optional admission control for saturated deployments, and corrects health reporting so that a container with a failed GPU worker no longer reports itself as ready.

Key Features#

  • Optional admission control and queue metrics. A saturated NIM can now shed load instead of queueing without bound. Set NIM_ENABLE_QUEUE_METRICS to true, together with NIM_MAX_CONCURRENT_REQUESTS, to enable it. Requests beyond the limit wait, and then receive 503 with a Retry-After header instead of blocking until an unrelated timeout expires.

    This NIM serves one request per GPU, so NIM_MAX_CONCURRENT_REQUESTS must equal the number of GPUs that the container can use. The container refuses to start when the two values disagree. A larger value admits requests that the GPU pool cannot run, and the queue depth reported on /v1/metrics then does not describe where those requests are waiting. Health and documentation endpoints bypass admission control. The /v1/models endpoint and the predict endpoint do not. Refer to Configuring the Boltz-2 NIM.

    While admission control is enabled, the following Prometheus series are published on /v1/metrics. The names match the vLLM-based NIMs, so autoscaling thresholds transfer between products.

    Metric

    Description

    nim_request_queue_depth

    Number of requests waiting for a GPU.

    nim_requests_running

    Number of requests running on a GPU.

    nim_request_queue_time_seconds

    Histogram of time spent waiting for a GPU.

    nim_requests_shed_total

    Requests rejected with 503, labeled queue_full or timeout.

  • nimlib upgrades from 0.19.5 to 0.22.0, which provides the admission control middleware and the queue metrics.

Fixed Issues#

  • A NIM with a failed GPU worker no longer reports itself as ready. When a GPU worker failed while loading models, /v1/health/ready continued to return true, and each request then blocked for 600 seconds before failing. Startup now fails with the error reported by the worker, and the container exits. Readiness continues to track worker liveness afterward, so a worker lost later to a CUDA fault or an out-of-memory condition also removes the NIM from rotation. A request that reaches a failed worker now fails in about five seconds instead of 600.

  • A fatal startup error no longer leaves the container running. Fatal conditions raised while the inference interface is built, such as a missing checkpoint under NIM_FAIL_WITHOUT_MODEL or an admission control limit that disagrees with the GPU count, previously left the process alive with no server listening. The container now exits with a non-zero status and reports the reason on standard error.

  • Addresses CVE findings across container dependencies.

Notes#

  • This release does not change prediction output. Response fields, their types, and predicted structures are unchanged.

  • Admission control is disabled unless you set NIM_ENABLE_QUEUE_METRICS. When it is unset, request handling and timeouts behave as they do in 1.9.0, and none of the queue metric series are registered. An uninstrumented NIM does not advertise metric names that it never populates.

Release 1.9.0#

Summary#

This release migrates inference onto a unified optimized backend, removes the TensorRT engine build and runtime path in favor of checkpoint-only assets on a PyTorch 26.05 base, routes structure prediction through a serial optimized pipeline, improves how prediction artifacts and custom checkpoints are handled, and increases inference performance and supported sequence lengths across GPUs.

Key Features#

  • Optimized backend: Structure and affinity inference now run on a single optimized backend (torch / CUDA-graph path) instead of a separate TensorRT-engine packaging path.

  • Checkpoint-only model assets: The model manifest is simplified to checkpoints and supporting assets (boltz2_conf.ckpt, boltz2_aff.ckpt, ccd.pkl, mols.tar). TensorRT engine build tooling is removed from the container workflow.

  • Single model profile: The 1.9.0 manifest ships one untagged profile (af1349b9f5f9a1a6a0404dea36dcc9499bcb25c9adc112b7cc9a93cae41f3262) used on all supported GPUs. Per-GPU TensorRT profile hashes from 1.8.0 are no longer valid. Leave NIM_MODEL_PROFILE unset, or pin that 1.9.0 profile ID. Refer to Support Matrix.

  • PyTorch 26.05 base: The NIM is rebuilt on the NVIDIA PyTorch 26.05 base image.

  • Serial optimized pipeline: Structure prediction uses a serial optimized pipeline path (no Ray dependency in the NIM container).

  • New optimized kernel backend includes:

    • cuEquivariance 0.11.1

    • NIMTools 1.12.0

    • CUDA Graph acceleration for short sequences (≤512 tokens)

  • Performance increased 3.9–6x on H200 and about 1.7–3.1x on B300, and sequence length limits increased across supported GPUs:

    GPU family

    Previous max sequence length

    New max sequence length

    Relative to previous

    L40S / A6000 / RTX 6000 Ada

    1665

    2048

    123.0%

    A100 / H100

    2560

    4096

    160%

    H200

    2560

    5120

    200%

    B200

    3072

    6144

    200%

    B300

    3072

    6144

    200%

    Actual maximum supported sequence length may exceed the values shown above, depending on available GPU VRAM.

    Benchmarks used the following predict configuration:

    {
      "recycling_steps": 10,
      "sampling_steps": 200,
      "diffusion_samples": 25,
      "max_parallel_samples": 25,
      "step_scale": 1.638,
      "output_format": "mmcif",
      "concatenate_msas": false,
      "write_full_pae": true,
      "write_full_pde": true
    }
    

    For large complexes (sequence length ≥5000), start with max_parallel_samples of 5. This field is the number of diffusion samples processed at the same time. Lower values reduce peak GPU memory; if an out-of-memory error occurs, reduce the value toward 1.

  • PAE/PDE artifact files: When write_full_pae or write_full_pde is enabled, full matrices are written as uncompressed .npz files under $NIM_OUTPUT_PATH/prediction_*/pae/ and $NIM_OUTPUT_PATH/prediction_*/pde/. The pae and pde fields in the JSON response are deprecated and remain null. Aggregate PDE scores (complex_pde_scores, complex_ipde_scores) continue to be returned in the JSON response.

  • Custom / fine-tuned checkpoints: Load same-architecture Boltz-2 weights from a mounted directory using MODEL_PATH, with optional filename overrides NIM_BOLTZ_CONF_CKPT_FILE and NIM_BOLTZ_AFFINITY_CKPT_FILE. Refer to Custom Checkpoints.

  • max_parallel_samples: New optional predict request field (range 1–25). Leave unset to use the optimized backend default. Lower values reduce GPU memory usage at the cost of longer runtime. For large cases (sequence length ≥5000), start with 5 and reduce toward 1 if an out-of-memory error occurs.

  • NIM_MAX_MSA_SEQS: New environment variable to cap MSA sequences retained when parsing request MSAs. Default is 4096; raise it (for example to 8192) to allow deeper MSAs.

  • CUDA memory headroom for confidence: The runtime frees CUDA cache before the confidence module to improve headroom on large complexes.

  • Security: Addressed CVE findings across container dependencies.

Removed Configuration#

The following environment variables are no longer supported and have been removed:

  • NIM_REQUIRE_TRT / NIM_FAIL_ON_NO_TRT

  • NIM_BUILD_TRT_ENGINES

  • NIM_BOLTZ_USE_TRT_ENGINES

Notes#

  • Mount a host directory at /opt/nim/output (or set NIM_OUTPUT_PATH) when enabling NIM_EXPOSE_CONFIDENCE_SCORES or write_full_pae / write_full_pde, so prediction artifacts land outside the model cache.

  • For best performance, run a warm-up pass with representative sequence lengths before measuring latency or throughput. The first request in a new length range can be slower because of runtime specialization (for example CUDA-graph capture on sequences ≤512 tokens).

Release 1.8.0#

Summary#

This release replaces the separate TensorRT and PyTorch backend options with a unified optimized backend built on new custom kernels, improves inference performance and supported sequence lengths, and adds affinity embedding outputs for binding affinity prediction.

Key Features#

  • Unified optimized backend: Removed the previous separate TensorRT (trt) and PyTorch (torch) backend selection. The NIM now uses a single new optimized backend that automatically selects the best execution path for each request.

  • Automatic backend selection:

    • Fully supports the Boltz2 OSS model

    • Short sequences (≤768 residues): TensorRT engines

    • Medium and long sequences (>768 residues): PyTorch backend with new optimized kernels

  • New optimized kernel backend includes:

  • Fused Adaptive Layer Norm

  • DualGemm sm80 (two variants)

  • Gated sigmoid sm80

  • FAv2/v3 TriangleAttent/AttentionPairBias

  • Pre-computed masks, with masks left-aligned into new kernels for dual GEMM and attention operations

  • Pre-computed buffers for Pairformer/DiT

  • Performance increased up to 1.7x on H100, and sequence length limits increased across supported GPUs:

    GPU family

    Previous max sequence length

    New max sequence length

    Relative to previous

    A100 / H100

    2048

    2560

    125%

    L40S / A6000 / RTX 6000 Ada

    1536

    1665

    108.4%

    B200 / B300 (Blackwell)

    2048

    3072

    150%

    Actual maximum supported sequence length may exceed the values shown above, depending on available GPU VRAM.

  • Affinity embedding outputs: Boltz2Affinity now exports embedding vectors as outputs. The new ligand-level request flag output_affinity_embedding defaults to false and requires predict_affinity=true on the same ligand. When enabled, the ligand’s entry under affinities includes affinity_embedding, model_1_affinity_embedding, and model_2_affinity_embedding (2D number arrays or null). Each embedding vector has 384 dimensions.

Removed Configuration#

The following environment variables are no longer supported and have been removed:

  • NIM_BOLTZ_ENABLE_DIFFUSION_TF32

  • NIM_BOLTZ_STRUCTURE_OPTIMIZED_BACKEND

  • NIM_BOLTZ_AFFINITY_OPTIMIZED_BACKEND

Notes#

  • For best performance with the unified optimized backend, run a warm-up pass with representative test-case sequences before measuring latency/throughput.

  • The first request in a given sequence-length range can be slower because of JIT compilation and runtime kernel specialization. As a result, you may observe a one-time startup latency when entering a new range (for example, short, medium, or long sequences). After that initial compilation for the range, subsequent requests in the same range should show improved performance.

Release 1.7.0#

Summary#

This release extends GPU support with NVIDIA Blackwell Ultra data center GPUs and Grace Blackwell Ultra superchips: NVIDIA B300 (NVIDIA-B300-SXM6-AC) and NVIDIA GB300 (Grace Blackwell Ultra) platforms.

Key Features#

  • NVIDIA B300 (NVIDIA-B300-SXM6-AC): Added support for this B300 Tensor Core GPU SKU with 288GB HBM3e memory per GPU (including DGX B300-class deployments)

  • NVIDIA GB300: Added support for GB300 Grace Blackwell Ultra superchips and rack-scale configurations built on Blackwell Ultra GPUs with high-capacity HBM3e (for example, GB300 NVL72 and DGX GB300)

Release 1.6.0#

Summary#

This release extends GPU support with the addition of NVIDIA GH200 superchips and the RTX PRO 6000 Blackwell Workstation Edition, adds Slurm support for HPC deployments, and adds support for NVIDIA H200.

Key Features#

  • NVIDIA H200: Added support for H200 with 141GB memory

  • NVIDIA GH200 144GB: Added support for GH200 with 144GB memory

  • NVIDIA RTX PRO 6000 Blackwell Workstation Edition: Added support for the professional workstation GPU with 96GB GDDR7 memory

  • Slurm support: Added support for running on Slurm-based HPC clusters (refer to Running on Slurm With Enroot in Getting Started)

Release 1.5.0#

Summary#

This release extends GPU support with the addition of the GB10 architecture, updates cuEquivariance for enhanced performance, and introduces new parameter configurations, template structure support, and output options.

Key Features#

  • GB10 DGX Spark Support: Added support for GB10 DGX Spark SKUs with sequence lengths up to 1536 residues

  • cuEquivariance integration: Updated to cuEquivariance 0.8.1 for GB10 support

  • Enhanced parameter ranges: Added matching parameter range for recycling_steps to 10 and diffusion_samples to 25, aligning with Boltz2 public model

  • PAE and PDE Matrix output: Support for returning the full PAE (Predicted Aligned Error) matrix using the write_full_pae = true flag and the full PDE (Predicted Distance Error) matrix using the write_full_pde = true flag in requests. Aggregate PDE scores (complex_pde_scores, complex_ipde_scores) are always included in responses.

  • Template structure support: Added support for template structures in both CIF and PDB formats, enabling users to provide structural templates for enhanced prediction accuracy

Release 1.4.0#

Summary#

This release addresses security vulnerabilities, integrates the latest cuEquivariance library, and introduces new features including telemetry control and confidence score persistence capabilities.

Key Features#

  • Security: Addressed all CVE (Common Vulnerabilities and Exposures) issues

  • cuEquivariance integration: Updated to cuEquivariance 0.7.0 for improved equivariant operations

  • Enhanced B200 support: Enabled trimul kernel optimization for B200 SKUs in PyTorch backend

  • Extended sequence length: Support for longer sequences up to 2048 residues on A100, H100, and B200 GPUs with inference optimizations

  • Confidence score persistence: Implemented Docker volume mounting flag to expose output directories with per-sample confidence JSON files (confidence_*_model_*.json), enabling result persistence in the NIM cache folder. These files include aggregate scores such as complex_pde and complex_ipde, but not full PAE/PDE matrices. Configure using:

    • Set NIM_EXPOSE_CONFIDENCE_SCORES=true to enable confidence score output

    • Set NIM_EXPOSE_CONFIDENCE_SCORES=false to disable (default)

    • To return full PAE/PDE matrices, use write_full_pae and write_full_pde in predict requests instead

  • Telemetry control: NIM Telemetry helps NVIDIA deliver a faster, more reliable experience with greater compatibility across a wide range of environments, while maintaining strict privacy protections and giving users full control.

    Benefits:

    • Enhances performance and reliability: Provides anonymous system and NIM-level insights that help NVIDIA identify bottlenecks, tune performance across hardware configurations, and improve runtime stability.

    • Improves compatibility across deployments: Helps detect and resolve version, driver, and environment compatibility issues early, reducing friction across diverse infrastructure setups.

    • Accelerates troubleshooting and bug resolution: Allows NVIDIA to diagnose errors and regressions faster, leading to quicker support response times and higher overall availability.

    • Informs smarter optimizations and future releases: Real-world, aggregated telemetry data helps guide the optimization of NIM runtimes, model packaging, and deployment workflows, ensuring updates target the scenarios that matter most to users.

    • Protects user privacy and data security: Collects only minimal, anonymous metadata, such as hardware type and NIM version. No user data, input sequences, or prediction results are collected.

    • Fully optional and configurable: Telemetry uses NIM_TELEMETRY_MODE (default 0, disabled). When set to 1, baseline collection (hardware and NIM metadata) can run on supported deployments. On RTX GPUs, collection stays off unless you also set NIM_TELEMETRY_ENABLE_ON_RTX=true (only applies when NIM_TELEMETRY_MODE is not 0).

    Configuration:

    • Set NIM_TELEMETRY_MODE=0 to disable telemetry (default)

    • Set NIM_TELEMETRY_MODE=1 to enable baseline telemetry where supported

    • Set NIM_TELEMETRY_ENABLE_ON_RTX=true to allow collection on RTX when NIM_TELEMETRY_MODE is not 0 (default for this variable is false)

    • Optional: NIM_TELEMETRY_ENABLE_LOGGING controls logging for telemetry operations when telemetry is enabled

    For more information about data privacy, what is collected, and how to configure telemetry, refer to:

Release 1.3.0#

Summary#

This release adds support for GB200 GPU with ARM architecture and extends sequence processing capabilities for both PyTorch and TensorRT backends.

Key Features#

  • GB200 Support: Added support for GB200 GPU with ARM architecture

  • Extended Sequence Processing:

    • PyTorch backend supports up to 4096 length sequences

    • TensorRT backend supports up to 2048 length sequences

  • Enhanced compatibility with ARM-based systems

Release 1.2.0#

Summary#

This release introduces support for longer sequence processing and additional GPU compatibility, along with performance optimizations.

Key Features#

  • Support longer sequence (2048 length) processing for TRT engines on A100, H100, and B200 with memory larger than 48GB

  • Support PyTorch and TRT engines with NVIDIA RTX6000 (RTX6000 and RTX6000-Ada GPU Workstation Edition GPU)

  • Improve model execution speed

  • Optimize memory usage during inference

Notes and Limitations#

  • Ensure you use this NIM with GPUs with at least 48 GB of VRAM.

  • TensorRT Engine Speedup: To achieve optimal prediction speedup using TRT engines, this version utilizes TensorRT 10.11.0. Refer to the TensorRT 10.11.0 Support Matrix to ensure your CUDA Driver and CUDA Toolkit versions are compatible.

    CUDA Requirements for TensorRT 10.11.0:

    • CUDA Driver: 12.9, 12.8 update 1, 12.6 update 3, 12.5 update 1, 12.4 update 1, 12.3 update 2, 12.2 update 2, 12.1 update 1, 12.0 update 1, 11.8, 11.7 update 1, 11.6 update 2, 11.5 update 2, 11.4 update 4, 11.3 update 1, 11.2 update 2, 11.1 update 1, 11.0 update 3

    • CUDA Toolkit: Must be compatible with the chosen CUDA Driver version

    • Platform Support: Linux x86-64, Windows x64, Linux SBSA, and NVIDIA JetPack

Note: While there are many options for tuning this NIM’s performance, for most users, the defaults will provide a balanced performance experience.

Release 1.1.0#

Summary#

Boltz-2 enables accurate biomolecular structure prediction from input sequences including proteins, DNA, RNA, and ligands. It also supports constrained guidance for complex structural prediction using bond and pocket specifications.

This release adds binding affinity prediction for ligands and complexes. The underlying Boltz-2 model has been updated to version 2.2.0. More Boltz-2 scores are included in the output, including aggregate PDE metrics (complex_pde_scores, complex_ipde_scores). There are minor bug fixes to improve the overall user experience.

Notes and Limitations#

Ensure you use this NIM with GPUs with at least 48 GB of VRAM.

Note: While there are many options for tuning this NIM’s performance, for most users the defaults will provide a balanced performance experience.