Configuring the Boltz-2 NIM#

The Boltz-2 NIM uses Docker containers. Each NIM has its own Docker container and there are several ways to configure it. The section below describes how to configure a NIM container.

GPU Selection#

By default, Docker can use all available GPUs on the system when it starts with the NVIDIA Container Runtime:

docker run --runtime=nvidia ...

In environments with a combination of GPUs, you can only expose specific GPUs inside the container using either:

  • The --gpus flag. For example, docker run --gpus='"device=1"' ...

  • The environment variable NVIDIA_VISIBLE_DEVICES. For example, to expose only Device 1, pass -e NVIDIA_VISIBLE_DEVICES=1. To expose GPU IDs 1 and 4, pass-e NVIDIA_VISIBLE_DEVICES=1,4.

The device IDs to use as inputs are listed in the output of nvidia-smi -L:

GPU 0: Tesla H100 80GB HBM3 (UUID: GPU-aaaaaaaa-bbbb-cccc-dddd-111111111111)
GPU 1: Tesla H100 80GB HBM3 (UUID: GPU-eeeeeeee-ffff-0000-1111-222222222222)

Refer to the NVIDIA Container Toolkit documentation for more instructions.

Environment Variables#

The Boltz-2 NIM container image on NVIDIA Container Registry is nvcr.io/nim/mit/boltz2:<version> (NGC org nim, team mit). Examples on this page use the substituted form nvcr.io/nim/mit/boltz2:1.10.0, which resolves to that path.

The following table describes the environment variables that can be passed into a NIM as a -e argument added to a docker run command:

ENV

Required?

Default

Notes

NGC_API_KEY

Yes

None

You must set this variable to the value of your personal NGC API key.

NIM_CACHE_PATH

No

/opt/nim/.cache

Location (in container) where the container caches model artifacts.

NIM_HTTP_API_PORT

No

8000

Publish the NIM service to the prescribed port inside the container. Make sure to adjust the port passed to the -p/--publish flag of Docker run to reflect that (ex: -p $NIM_HTTP_API_PORT:$NIM_HTTP_API_PORT). The left-hand side of this : is your host address:port, and does NOT have to match with $NIM_HTTP_API_PORT. The right-hand side of the : is the port inside the container and MUST match the port the NIM listens on: use the value of NIM_HTTP_API_PORT when set, otherwise 8000 (the default in this column—not a different port). Supported endpoints are /v1/license (Returns the license information), /v1/metadata (Returns metadata including asset information, license information, model information, and version) and /v1/metrics (Exposes Prometheus metrics using an ASGI app endpoint).

NIM_LOG

No

INFO

Controls NIM service logging. Available options are DEBUG, INFO, WARNING, ERROR, and CRITICAL.

NIM_LOG_LEVEL

No

INFO

Alternative NIM logging level control. Available options are DEBUG, INFO, WARNING, ERROR, and CRITICAL.

MODEL_PATH

No

Unset

Hard override of the NIM model root. When set, checkpoints and supporting assets are loaded from this directory instead of the NGC cache. Used for custom / fine-tuned checkpoints.

NIM_DISABLE_MODEL_DOWNLOAD

No

Unset

When set to 1/true, skip downloading model artifacts from NGC. Use together with MODEL_PATH when serving local checkpoints.

NIM_BOLTZ_CONF_CKPT_FILE

No

boltz2_conf.ckpt

Filename of the structure checkpoint inside MODEL_PATH or the NGC cache root.

NIM_BOLTZ_AFFINITY_CKPT_FILE

No

boltz2_aff.ckpt

Filename of the affinity checkpoint inside MODEL_PATH or the NGC cache root.

NIM_DEFAULT_RANDOM_SEED

No

42

Controls the random seed used for inference.

NIM_TELEMETRY_MODE

No

0

String 0 disables telemetry (default). 1 enables baseline collection (hardware and NIM information). On RTX GPUs, collection stays off unless you also set NIM_TELEMETRY_ENABLE_ON_RTX=true. Refer to NIM Telemetry. For privacy and full configuration details, refer to NVIDIA’s Privacy Policy and NIM Telemetry Settings.

NIM_TELEMETRY_ENABLE_ON_RTX

No

false

String true allows telemetry on RTX GPUs when NIM_TELEMETRY_MODE is not 0. Default false. Has no effect when telemetry is disabled. Refer to NIM Telemetry.

NIM_TELEMETRY_ENABLE_LOGGING

No

true

Enables logging for telemetry operations when set to true. Only applicable when NIM_TELEMETRY_MODE=1.

NIM_OUTPUT_PATH

No

/opt/nim/output

Root directory for exposed prediction artifacts (prediction_*/{confidence_scores,pae,pde}/). Falls back to NIM_CACHE_PATH when unset. Mount a host directory here when persisting confidence JSONs or full PAE/PDE .npz files.

NIM_EXPOSE_CONFIDENCE_SCORES

No

false

Controls whether to expose per-sample confidence JSON files (confidence_*_model_*.json) under $NIM_OUTPUT_PATH/prediction_*/confidence_scores/. Set to true to persist files (including scalar metrics such as complex_pde and complex_ipde), set to false to disable (default). This setting does not write full PAE or PDE matrices; use write_full_pae and write_full_pde in predict requests to write full matrices as .npz files under $NIM_OUTPUT_PATH.

NIM_MAX_MSA_SEQS

No

4096

Cap on MSA sequences retained when parsing request MSAs. Raise this to allow deeper MSAs (for example 8192).

NIM_MAX_POLYMER_INPUTS

No

12

Maximum number of polymer chains allowed in a single predict request.

NIM_MAX_LIGAND_INPUTS

No

20

Maximum number of ligands allowed in a single predict request.

NIM_MAX_POLYMER_LENGTH

No

4096

Maximum length for an individual polymer sequence.

NIM_MAX_TEMPLATES_PER_POLYMER

No

4

Maximum number of structural templates allowed per protein polymer.

NIM_ENABLE_QUEUE_METRICS

No

false

Set to true to enable admission control and publish the request queue metrics on /v1/metrics. Requires NIM_MAX_CONCURRENT_REQUESTS. Refer to Admission Control.

NIM_MAX_CONCURRENT_REQUESTS

No

Unset

Number of requests served at the same time. This NIM runs one request per GPU, so set this variable to the number of GPUs that the container can use. Required when NIM_ENABLE_QUEUE_METRICS is enabled. The container refuses to start when the two values disagree.

NIM_MAX_QUEUE_DEPTH

No

Unset

Number of requests allowed to wait for a free GPU. Beyond this number, a request is rejected on arrival with 503 instead of joining a queue that it is unlikely to clear. Both an unset value and 0 mean unbounded. Set the value to 1 for the smallest possible queue.

NIM_QUEUE_TIMEOUT_SECONDS

No

300

How long a request waits for a free GPU before it is rejected with 503. A value of 0 waits indefinitely.

NIM_QUEUE_RETRY_AFTER_SECONDS

No

30

Value sent in the Retry-After header of a rejected request. The value is jittered so that a rejected batch of requests does not return at the same moment. Raise it to match the service time that you observe. Boltz-2 predictions commonly take minutes, and a much lower value invites clients to retry against a NIM that cannot serve them yet.

NIM_GPU_ACQUIRE_TIMEOUT_SECONDS

No

Derived

How long a request waits internally for a GPU. Leave this variable unset. The value is derived so that it is longer than NIM_QUEUE_TIMEOUT_SECONDS, which lets admission control reject the request first and return 503 with a Retry-After header. A value of 0 waits indefinitely.

For additional Boltz-2 tuning variables such as NIM_MODEL_PROFILE, refer to Optimization.

Admission Control#

By default, the NIM accepts every request, and each request waits its turn for a GPU. In a saturated deployment that queue is unbounded, and a client cannot distinguish a slow prediction from one that never starts.

Setting NIM_ENABLE_QUEUE_METRICS to true bounds the queue. Requests beyond NIM_MAX_CONCURRENT_REQUESTS wait, and are rejected with 503 and a Retry-After header after NIM_QUEUE_TIMEOUT_SECONDS elapses, or when NIM_MAX_QUEUE_DEPTH is exceeded. While admission control is enabled, /v1/metrics publishes nim_request_queue_depth, nim_requests_running, nim_request_queue_time_seconds, and nim_requests_shed_total.

docker run -it --rm --name=boltz2 \
    --runtime=nvidia \
    --gpus 4 \
    -e NGC_API_KEY \
    -e NIM_ENABLE_QUEUE_METRICS=true \
    -e NIM_MAX_CONCURRENT_REQUESTS=4 \
    -e NIM_MAX_QUEUE_DEPTH=8 \
    -e NIM_QUEUE_TIMEOUT_SECONDS=600 \
    -e NIM_QUEUE_RETRY_AFTER_SECONDS=180 \
    -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
    -p 8000:8000 \
    nvcr.io/nim/mit/boltz2:1.10.0

The following table describes the three requirements that admission control cannot infer for you.

Requirement

Reason

NIM_MAX_CONCURRENT_REQUESTS must equal the number of GPUs that the container can use.

The NIM serves one request per GPU, and its worker pool enforces that limit regardless of this setting. A larger value moves the queue to a place that the metrics cannot observe, and a smaller value leaves GPUs idle. The container refuses to start when the two values disagree, rather than report a limit that does not hold.

Keep NIM_HTTP_MAX_WORKERS at its default value of 1.

Each worker process enforces the concurrency limit separately, so several workers admit a multiple of the intended number of requests.

Health and documentation endpoints bypass admission control.

A saturated NIM continues to answer /v1/health/ready and to serve /docs. The /v1/models endpoint and the predict endpoint are subject to admission control.

NIM Telemetry#

NIM Telemetry collects minimal host and NIM information to help improve performance, reliability, and compatibility across deployments. Collection is disabled by default on RTX GPUs and can be enabled explicitly when you opt in.

NIM_TELEMETRY_MODE#

Controls telemetry collection mode.

  • 0 — Disables telemetry.

  • 1 — Enables baseline collection (hardware and NIM information).

Default: 0

NIM_TELEMETRY_ENABLE_ON_RTX#

Set to true to allow telemetry collection on RTX GPUs. This setting applies only when NIM_TELEMETRY_MODE is not 0 (for example, when baseline collection is enabled on the deployment).

Default: false

For what is collected, retention, and advanced options, refer to NIM Telemetry Settings and NVIDIA’s Privacy Policy.

Volumes#

The following table describes the paths inside the container into which local paths can be mounted.

Container path

Required

Notes

Docker argument example

/opt/nim/.cache (or NIM_CACHE_PATH if present)

Not required, but if this volume is not mounted, the container will do a fresh download of the model each time it is brought up.

This is the directory within which models are downloaded inside the container. It is very important that this directory can be accessed from inside the container. This can be achieved by setting the permissions of the local directory to read-write-execute (777). For example, to use ~/.cache/nim as the host machine directory for caching models, first do mkdir -p ~/.cache/nim, then chmod 777 ~/.cache/nim before running the docker run command.

-v ~/.cache/nim:/opt/nim/.cache

/opt/nim/output (or NIM_OUTPUT_PATH if present)

Not required

Destination for exposed prediction artifacts such as confidence JSONs and full PAE/PDE .npz files under prediction_*/. Keep this separate from the model cache so large matrices do not fill NIM_CACHE_PATH.

-v ./output:/opt/nim/output -e NIM_OUTPUT_PATH=/opt/nim/output

Logging Configuration#

The Boltz-2 NIM provides logging environment variables to control verbosity for different components. You can configure these when starting the container to adjust the level of detail in logs for debugging or production use.

Example: Running with Logging Configuration#

The following example shows how to run the NIM with standard logging configuration:

export LOCAL_NIM_CACHE=~/.cache/nim
export NGC_API_KEY=<Your NGC API Key>

docker run --rm --name boltz2 --runtime=nvidia \
    --shm-size=16G \
    -e NGC_API_KEY \
    -e NIM_LOG=INFO \
    -e NIM_LOG_LEVEL=INFO \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 8000:8000 \
    nvcr.io/nim/mit/boltz2:1.10.0

Example: Debug Logging for Troubleshooting#

To enable verbose logging for troubleshooting issues, set the log levels to their most detailed settings:

export LOCAL_NIM_CACHE=~/.cache/nim
export NGC_API_KEY=<Your NGC API Key>

docker run --rm --name boltz2 --runtime=nvidia \
    --shm-size=16G \
    -e NGC_API_KEY \
    -e NIM_LOG=DEBUG \
    -e NIM_LOG_LEVEL=DEBUG \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 8000:8000 \
    nvcr.io/nim/mit/boltz2:1.10.0