Optimization with Boltz2-NIM#
This section details the options available for optimizing the Boltz-2 NIM. Note that achieving optimal performance depends on many factors and may require settings unique to your individual deployment.
Note
For most users, the default settings of the NIM will provide good performance that balances the throughput, latency, resource utilization, and complexity of using the NIM. We recommend only changing these options after consulting with an expert about your specific use case and any unique performance requirements it may have.
Automatic Profile Selection#
The Boltz-2 NIM selects a model profile at startup for download and caching. Starting in release 1.9.0, the shipped manifest contains a single untagged profile. Auto-selection uses that profile on every supported GPU. Inference always runs on a unified optimized backend with a torch / CUDA-graph path. There is no separate TensorRT-engine backend to select per request or per GPU.
Selecting a Profile Manually#
Leave NIM_MODEL_PROFILE unset for automatic selection. To pin the shipped 1.9.0 profile for reproducibility, set NIM_MODEL_PROFILE to the profile ID in the Support Matrix. Using a hash from a previous release (for example a 1.8.0 per-GPU TensorRT profile) fails because those IDs are not in the 1.9.0 manifest.
# Optional: pin the 1.9.0 unified profile
export NIM_MODEL_PROFILE=af1349b9f5f9a1a6a0404dea36dcc9499bcb25c9adc112b7cc9a93cae41f3262
docker run --rm --name boltz2 --runtime=nvidia \
--shm-size=16G \
-e NGC_API_KEY \
-e NIM_MODEL_PROFILE \
-v $LOCAL_NIM_CACHE:/opt/nim/.cache \
-p 8000:8000 \
nvcr.io/nim/mit/boltz2:1.10.0
Optimized Backend#
Starting in release 1.9.0, the Boltz-2 NIM uses a unified optimized backend for structure and affinity inference:
Checkpoint-only assets (
boltz2_conf.ckpt,boltz2_aff.ckpt, plusccd.pkl/mols.tar)Torch / CUDA-graph optimized execution (no TensorRT engine build or load path)
CUDA Graph acceleration for short sequences (≤512 tokens)
Serial optimized pipeline path inside the NIM (no Ray dependency)
cuEquivariance 0.11.1 and NIMTools 1.12.0
Release 1.8.0 introduced a unified optimized backend with custom kernels and affinity embedding outputs. Release 1.9.0 keeps those capabilities while replacing the TensorRT-engine packaging path with the optimized backend described above. Refer to Release 1.9.0 and Release 1.8.0.
For most users, no backend configuration is required.
Deploying the NIM on a Multi-GPU System#
The Boltz-2 NIM is designed to run with one or more NVIDIA GPUs. When increasing the number of GPUs allocated to the NIM, it is recommended to also increase the allocated number of CPU cores and RAM. As a rule of thumb, for each additional GPU allocated, you should also allocate another additional 12 CPU cores and 32 GB of additional system RAM.
Adjusting Start-Time NIM Input Limits#
The Boltz-2 NIM can be configured at startup time using environment variables to control input limits and resource usage. These settings help prevent excessively large requests that could impact performance or cause out-of-memory errors.
Environment Variables#
NIM_MAX_POLYMER_INPUTS#
Default:
12Type: Integer
Description: Sets the maximum number of polymer chains (DNA, RNA, or protein) that can be included in a single prediction request.
NIM_MAX_LIGAND_INPUTS#
Default:
20Type: Integer
Description: Sets the maximum number of ligands that can be included in a single prediction request.
NIM_MAX_POLYMER_LENGTH#
Default:
4096Type: Integer
Description: Sets the maximum allowed length for individual polymer sequences (number of residues/nucleotides).
NIM_MAX_TEMPLATES_PER_POLYMER#
Default:
4Type: Integer
Description: Sets the maximum number of structural templates that can be provided for each protein polymer.
NIM_MAX_MSA_SEQS#
Default:
4096Type: Integer
Description: Cap on MSA sequences retained when parsing request MSAs. Raise this (for example to
8192) to allow deeper MSAs.
Usage#
Set these environment variables before starting the NIM to customize the input limits:
export NIM_MAX_POLYMER_INPUTS=12
export NIM_MAX_LIGAND_INPUTS=20
export NIM_MAX_POLYMER_LENGTH=4096
export NIM_MAX_TEMPLATES_PER_POLYMER=4
export NIM_MAX_MSA_SEQS=4096
# Start the NIM with custom limits
docker run --rm --name boltz2-nim --runtime=nvidia \
--shm-size=16G \
-e NGC_API_KEY \
-e NIM_MAX_POLYMER_INPUTS \
-e NIM_MAX_LIGAND_INPUTS \
-e NIM_MAX_POLYMER_LENGTH \
-e NIM_MAX_TEMPLATES_PER_POLYMER \
-e NIM_MAX_MSA_SEQS \
-v $LOCAL_NIM_CACHE:/opt/nim/.cache \
-p 8000:8000 \
nvcr.io/nim/mit/boltz2:1.10.0
Note
Increasing these values may result in runtime instability of the NIM, especially with regards to memory usage. Note that these limits are applied at startup and cannot be changed without restarting the NIM. Choose values that balance your performance requirements with the computational resources available to your deployment.
Optimization Parameters#
The Boltz-2 NIM provides several parameters that can be tuned to optimize performance for your specific use case:
Recycling Steps#
The recycling_steps parameter (range: 1-10, default: 3) controls the number of iterative refinement steps. Higher values generally improve accuracy but increase computation time.
Sampling Steps#
The sampling_steps parameter (range: 10-1,000, default: 50) controls the number of diffusion sampling steps. More steps can improve quality but significantly increase runtime.
Diffusion Samples#
The diffusion_samples parameter (range: 1-25, default: 1) controls how many independent structure predictions are generated. Multiple samples provide diversity but multiply the computational cost.
Max Parallel Samples#
The max_parallel_samples parameter (range: 1-25, optional) caps how many diffusion samples are processed concurrently. Leave unset to use the optimized backend default. Lower values reduce peak GPU memory usage but increase wall-clock runtime. For large complexes (sequence length ≥5000), start with 5 and reduce toward 1 if an out-of-memory error occurs.
Full PAE and PDE Matrices#
The write_full_pae and write_full_pde request flags write full error matrices as uncompressed .npz files under $NIM_OUTPUT_PATH/prediction_*/pae/ and $NIM_OUTPUT_PATH/prediction_*/pde/. Matrices are not embedded in the JSON response (pae / pde fields remain null). Each enabled flag writes float32 arrays of roughly O(tokens² × diffusion_samples) values; on-disk size is close to the raw array size because these NPZ archives are not compressed. For large complexes, enable these flags only when you need pairwise error analysis, and mount $NIM_OUTPUT_PATH to a host volume with enough free space.
Step Scale#
The step_scale parameter (range: 0.5-5.0, default: 1.638) controls the magnitude of change applied at each diffusion sampling step. Larger values can speed up convergence but may reduce coordinate precision; smaller values can improve precision but require more sampling steps. Recommended range: 1.0-2.0.
Note
These parameters offer a tradeoff between prediction quality and computational cost. For production workloads, consider starting with default values and adjusting based on your specific quality and latency requirements.
Binding Affinity Prediction Parameters#
When performing binding affinity prediction with ligands, additional parameters control the quality and performance of the affinity calculation:
Affinity Sampling Steps#
The sampling_steps_affinity parameter (range: 10-1,000, default: 200) controls the number of diffusion sampling steps specifically for binding affinity prediction. This is separate from and typically higher than the structural sampling steps to ensure accurate affinity estimates. Higher values may improve accuracy but increase runtime.
Affinity Diffusion Samples#
The diffusion_samples_affinity parameter (range: 1-10, default: 5) controls how many independent affinity predictions are generated and averaged. Multiple samples improve the reliability of affinity estimates but increase runtime. Higher values may improve reliability but significantly increase runtime.
Molecular Weight Correction#
The affinity_mw_correction parameter (default: False) enables molecular weight correction for affinity predictions, which can improve accuracy for certain types of ligands.
Note
Binding affinity prediction requires significantly more computational resources than structure prediction alone. The affinity-specific parameters typically use higher values than their structural counterparts to ensure accurate binding energy estimates.
Ligand-Specific Performance Considerations#
Working with ligands and binding affinity prediction has unique performance characteristics:
Memory scaling: Protein-ligand complexes require additional GPU memory compared to protein-only predictions
Runtime impact: Binding affinity prediction can increase total runtime by 2-4x compared to structure-only prediction
Concurrent requests: Due to higher memory usage, fewer concurrent requests may be possible when using binding affinity prediction
Ligand count limits: The NIM supports up to 20 ligands per request, but binding affinity can only be predicted for one ligand at a time
The following examples highlight how to use the various optimization parameters.
import requests
# Higher recycling_steps for improved accuracy
payload = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": "YOUR_PROTEIN_SEQUENCE_HERE"
}
],
"recycling_steps": 5 # Default is 3. Higher values may improve accuracy.
}
response = requests.post(
"http://localhost:8000/biology/mit/boltz2/predict",
json=payload
)
response.raise_for_status()
print(response.json())
import requests
# Higher sampling_steps for improved quality
payload = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": "YOUR_PROTEIN_SEQUENCE_HERE"
}
],
"sampling_steps": 100 # Default is 50. More steps can improve quality but increase runtime.
}
response = requests.post(
"http://localhost:8000/biology/mit/boltz2/predict",
json=payload
)
response.raise_for_status()
print(response.json())
import requests
# Generate multiple distinct structures
payload = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": "YOUR_PROTEIN_SEQUENCE_HERE"
}
],
"diffusion_samples": 3 # Default is 1. Generates 3 candidate structures.
}
response = requests.post(
"http://localhost:8000/biology/mit/boltz2/predict",
json=payload
)
response.raise_for_status()
print(response.json())
import requests
# Lower step_scale for more diversity among samples
payload = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": "YOUR_PROTEIN_SEQUENCE_HERE"
}
],
"step_scale": 1.5 # Default is 1.638. Lower values increase diversity.
}
response = requests.post(
"http://localhost:8000/biology/mit/boltz2/predict",
json=payload
)
response.raise_for_status()
print(response.json())
import requests
# Predict protein-ligand complex structure
payload = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": "YOUR_PROTEIN_SEQUENCE_HERE"
}
],
"ligands": [
{
"id": "L1",
"smiles": "CC(=O)OC1=CC=CC=C1C(=O)O" # Aspirin
}
],
"recycling_steps": 3,
"sampling_steps": 50
}
response = requests.post(
"http://localhost:8000/biology/mit/boltz2/predict",
json=payload
)
response.raise_for_status()
print(response.json())
import requests
# Predict hemoglobin-heme binding affinity
# Using human hemoglobin alpha chain sequence
hemoglobin_alpha = (
"MVLSPADKTNVKAAWGKVGAHAGEYGAEALERMFLSFPTTKTYFPHFDLSHGSAQVKGHGKKVADALT"
"NAVAHVDDMPNALSALSDLHAHKLRVDPVNFKLLSHCLLVTLAAHLPAEFTPAVHASLDKFLASVSTV"
"LTSKYR"
)
payload = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": hemoglobin_alpha
}
],
"ligands": [
{
"id": "HEME",
"smiles": "[Fe+2].C1=CC2=NC1=CC3=NC(=CC4=NC(=CC5=NC(=C2)C=C5)C=C4)C=C3",
"predict_affinity": True
}
],
"recycling_steps": 4, # Higher for complex binding
"sampling_steps": 75,
"sampling_steps_affinity": 250, # High accuracy for heme binding
"diffusion_samples_affinity": 7, # Multiple samples for reliable binding estimate
"affinity_mw_correction": True # Useful for metallic centers (e.g. heme iron)
}
response = requests.post(
"http://localhost:8000/biology/mit/boltz2/predict",
json=payload
)
response.raise_for_status()
result = response.json()
# Display both structure and affinity results
print(f"Hemoglobin-heme complex prediction completed")
print(f"Confidence score: {result['confidence_scores'][0]:.3f}")
if result.get("affinities"):
for ligand_id, affinity_data in result["affinities"].items():
print(f"\nBinding affinity results for {ligand_id}:")
if affinity_data.get("affinity_pic50"):
print(f" pIC50: {affinity_data['affinity_pic50'][0]:.2f}")
if affinity_data.get("affinity_pred_value"):
print(f" log(IC50): {affinity_data['affinity_pred_value'][0]:.2f}")
if affinity_data.get("affinity_probability_binary"):
print(f" Binding probability: {affinity_data['affinity_probability_binary'][0]:.3f}")
Querying the NIM Repeatedly#
In computational biology and drug discovery, it is common to need to analyze hundreds, thousands, or even millions of protein sequences to find candidates with the desired properties. This section details some useful patterns for users that want to analyze more than one input sequence.
Note
The Boltz-2 NIM has been optimized for a balance of throughput and latency, and the performance of repeated queries may not be in line with published benchmarks due to the complexity of scheduling such workloads in the NIM. Factors, such as the underlying hardware and software stack, number of concurrent users, and system load, may impact the latency and throughput of the NIM.
Running Repeated Queries Against the NIM Serially#
Repeated queries against the NIM can be submitted as subsequent requests. For example, requests from a list of sequences
can be submitted using a for loop. Each request is blocking, which means a maximum of one request from a single
submitter will run at a time when calling the NIM in this manner. The following is an example of submitting multiple requests to the NIM:
Note
In the following examples, we use small multiple sequence alignments to demonstrate the usage of the API. Actual multiple sequence alignments may be much larger.
import requests
import json
def main():
url = "http://localhost:8000/biology/mit/boltz2/predict"
proteins = [
{
"name": "Green Fluorescent Protein (GFP)",
"sequence": (
"MSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTFCYGD"
"QIQEQYKGIPLDGDQVQAVNGHEFEIEGEGEGRPYEGTQTAQ"
),
"msa": {
"uniref90": {
"a3m": {
"format": "a3m",
"alignment": (
">seq1\nMSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTFCYGD"
"QIQEQYKGIPLDGDQVQAVNGHEFEIEGEGEGRPYEGTQTAQ\n"
">seq2\nMSKGEELFTGVVPILVELDGDVNGHKFSVSGEGEGDATYGKLTLKFICTTGKLPVPWPTLVTTFCYGD"
"QIQEQYKGIPLDGDQVQAVNGHEFEIEGEGEGRPYEGTQTAQ"
)
}
}
}
},
{
"name": "Tumor Protein p53",
"sequence": (
"MEEPQSDPSVEPPLSQETFSDLWKLLPENNVLSPLPSQAMDDLMLSPDDIEQWFTEDPGPDEAPRMPEA"
"APPVAPAPAAPTPAAPAPAPSWPLSSSVPSQKTYQGSYGFRLGFLHSGTAKSVTCTYSPALNKMFCQLA"
"KTCPVQLWVDSTPPPGTRVRAMAIYKKSQHMTEVVRRCPHHERCSDSDGLAPPQHLIRVEGNLRVEYLD"
"DPSKYLQW"
),
"msa": {
"uniref90": {
"a3m": {
"format": "a3m",
"alignment": (
">seq1\nMEEPQSDPSVEPPLSQETFSDLWKLLPENNVLSPLPSQAMDDLMLSPDDIEQWFTEDPGPDEAPRMPEA"
"APPVAPAPAAPTPAAPAPAPSWPLSSSVPSQKTYQGSYGFRLGFLHSGTAKSVTCTYSPALNKMFCQLA"
"KTCPVQLWVDSTPPPGTRVRAMAIYKKSQHMTEVVRRCPHHERCSDSDGLAPPQHLIRVEGNLRVEYLD"
"DPSKYLQW\n"
">seq2\nMEEPQSDPSVEPPLSQETFSDLWKLLPENNVLSPLPSQAMDDLMLSPDDIEQWFTEDPGPDEAPRMPEA"
"APPVAPAPAAPTPAAPAPAPSWPLSSSVPSQKTYQGSYGFRLGFLHSGTAKSVTCTYSPALNKMFCQLA"
"KTCPVQLWVDSTPPPGTRVRAMAIYKKSQHMTEVVRRCPHHERCSDSDGLAPPQHLIRVEGNLRVEYLD"
"DPSKYLQW"
)
}
}
}
},
{
"name": "Lactose Operon Repressor (LacI)",
"sequence": (
"MKPVTLYDVAEYAGVSYQTVSRVVNQASHVSAKTREKVEAAMAELNYIPNRVAQQLAGKQNLKDGDPTR"
"ADKKSIEYSASVSRQQSYSIKKNLIDQFEAQKPSLTGMSADSQIGQVTKDAQAMIKAIGVNLLQFPRQ"
"SPGDLEQGVNLTPCTLNTVTQTSLSVRGDKLIAEIGDKVAASEN"
),
"msa": {
"uniref90": {
"a3m": {
"format": "a3m",
"alignment": (
">seq1\nMKPVTLYDVAEYAGVSYQTVSRVVNQASHVSAKTREKVEAAMAELNYIPNRVAQQLAGKQNLKDGDPTR"
"ADKKSIEYSASVSRQQSYSIKKNLIDQFEAQKPSLTGMSADSQIGQVTKDAQAMIKAIGVNLLQFPRQ"
"SPGDLEQGVNLTPCTLNTVTQTSLSVRGDKLIAEIGDKVAASEN\n"
">seq2\nMKPVTLYDVAEYAGVSYQTVSRVVNQASHVSAKTREKVEAAMAELNYIPNRVAQQLAGKQNLKDGDPTR"
"ADKKSIEYSASVSRQQSYSIKKNLIDQFEAQKPSLTGMSADSQIGQVTKDAQAMIKAIGVNLLQFPRQ"
"SPGDLEQGVNLTPCTLNTVTQTSLSVRGDKLIAEIGDKVAASEN"
)
}
}
}
},
{
"name": "Bovine Serum Albumin (BSA)",
"sequence": (
"MKWVTFISLLFLFSSAYSRGVFRRDTHKSEIAHRFKDLGEENFKALVLIAFAQYLQQCPFDEHVKLVNE"
"GTKPVETVTKLVTDLTKVHTECCHGDLLECADDRADLAKYICDNQDTISSKLKECCDKPLLEKSHCIAE"
"VFCKYKEHKEMPFPKCCETSLVNRRPCFSALTPDETYVPKAFDEKLFTFHADICTLPDTEKQIKKQTAL"
"VELVKHKPKATKEQLKAVMDDFAAFVEKCCKADDKETCFAEEGKKLVAASQAALGL"
),
"msa": {
"uniref90": {
"a3m": {
"format": "a3m",
"alignment": (
">seq1\nMKWVTFISLLFLFSSAYSRGVFRRDTHKSEIAHRFKDLGEENFKALVLIAFAQYLQQCPFDEHVKLVNE"
"GTKPVETVTKLVTDLTKVHTECCHGDLLECADDRADLAKYICDNQDTISSKLKECCDKPLLEKSHCIAE"
"VFCKYKEHKEMPFPKCCETSLVNRRPCFSALTPDETYVPKAFDEKLFTFHADICTLPDTEKQIKKQTAL"
"VELVKHKPKATKEQLKAVMDDFAAFVEKCCKADDKETCFAEEGKKLVAASQAALGL\n"
">seq2\nMKWVTFISLLFLFSSAYSRGVFRRDTHKSEIAHRFKDLGEENFKALVLIAFAQYLQQCPFDEHVKLVNE"
"GTKPVETVTKLVTDLTKVHTECCHGDLLECADDRADLAKYICDNQDTISSKLKECCDKPLLEKSHCIAE"
"VFCKYKEHKEMPFPKCCETSLVNRRPCFSALTPDETYVPKAFDEKLFTFHADICTLPDTEKQIKKQTAL"
"VELVKHKPKATKEQLKAVMDDFAAFVEKCCKADDKETCFAEEGKKLVAASQAALGL"
)
}
}
}
}
]
responses = []
# For each protein, submit the request and store the response
for protein in proteins:
data = {
"polymers": [
{
"id": "A",
"molecule_type": "protein",
"sequence": protein["sequence"],
"msa": protein["msa"]
}
],
"recycling_steps": 3,
"sampling_steps": 50,
"diffusion_samples": 1,
"step_scale": 1.638,
"output_format": "mmcif"
}
response = None # Initialize response
response_data = None # Initialize response_data
try:
# Use the 'json' parameter instead of 'data'
response = requests.post(url, json=data)
# Attempt to parse the JSON response
response_data = response.json()
if response.ok:
print(f"Structure prediction for {protein['name']} succeeded.")
else:
print(f"Structure prediction for {protein['name']} failed: {response.status_code} {response.text}")
except requests.exceptions.RequestException as req_err:
# Catch any request-related errors
print(f"Structure prediction for {protein['name']} failed: {req_err}")
except json.JSONDecodeError:
# Catch JSON parsing errors
print(f"Response from {protein['name']} could not be decoded as JSON.")
# Store the response along with the protein name and sequence
responses.append({
"protein": protein["name"],
"sequence": protein["sequence"],
"response": response_data,
"status_code": response.status_code if response else None,
"text": response.text if response else None
})
# Print the responses
for res in responses:
print(f"Protein: {res['protein']}")
print(f"Status Code: {res['status_code']}")
if res['response']:
print("Response Data:", json.dumps(res['response'], indent=2))
else:
print("No response data available.")
print("-" * 40)
if __name__ == "__main__":
main()