Configuring the Boltz-2 NIM#
The Boltz-2 NIM uses Docker containers. Each NIM has its own Docker container and there are several ways to configure it. The section below describes how to configure a NIM container.
GPU Selection#
By default, Docker can use all available GPUs on the system when it starts with the NVIDIA Container Runtime:
docker run --runtime=nvidia ...
In environments with a combination of GPUs, you can only expose specific GPUs inside the container using either:
The
--gpusflag. For example,docker run --gpus='"device=1"' ...The environment variable
NVIDIA_VISIBLE_DEVICES. For example, to expose only Device 1, pass-e NVIDIA_VISIBLE_DEVICES=1. To expose GPU IDs 1 and 4, pass-e NVIDIA_VISIBLE_DEVICES=1,4.
The device IDs to use as inputs are listed in the output of nvidia-smi -L:
GPU 0: Tesla H100 80GB HBM3 (UUID: GPU-aaaaaaaa-bbbb-cccc-dddd-111111111111)
GPU 1: Tesla H100 80GB HBM3 (UUID: GPU-eeeeeeee-ffff-0000-1111-222222222222)
Refer to the NVIDIA Container Toolkit documentation for more instructions.
Environment Variables#
The Boltz-2 NIM container image on NVIDIA Container Registry is nvcr.io/nim/mit/boltz2:<version> (NGC org nim, team mit). Examples on this page use the substituted form nvcr.io/nim/mit/boltz2:1.10.0, which resolves to that path.
The following table describes the environment variables that can be passed into a NIM as a -e argument added to a docker run command:
ENV |
Required? |
Default |
Notes |
|---|---|---|---|
|
Yes |
None |
You must set this variable to the value of your personal NGC API key. |
|
No |
|
Location (in container) where the container caches model artifacts. |
|
No |
|
Publish the NIM service to the prescribed port inside the container. Make sure to adjust the port passed to the |
|
No |
|
Controls NIM service logging. Available options are |
|
No |
|
Alternative NIM logging level control. Available options are |
|
No |
Unset |
Hard override of the NIM model root. When set, checkpoints and supporting assets are loaded from this directory instead of the NGC cache. Used for custom / fine-tuned checkpoints. |
|
No |
Unset |
When set to |
|
No |
|
Filename of the structure checkpoint inside |
|
No |
|
Filename of the affinity checkpoint inside |
|
No |
|
Controls the random seed used for inference. |
|
No |
|
String |
|
No |
|
String |
|
No |
|
Enables logging for telemetry operations when set to |
|
No |
|
Root directory for exposed prediction artifacts ( |
|
No |
|
Controls whether to expose per-sample confidence JSON files ( |
|
No |
|
Cap on MSA sequences retained when parsing request MSAs. Raise this to allow deeper MSAs (for example |
|
No |
|
Maximum number of polymer chains allowed in a single predict request. |
|
No |
|
Maximum number of ligands allowed in a single predict request. |
|
No |
|
Maximum length for an individual polymer sequence. |
|
No |
|
Maximum number of structural templates allowed per protein polymer. |
|
No |
|
Set to |
|
No |
Unset |
Number of requests served at the same time. This NIM runs one request per GPU, so set this variable to the number of GPUs that the container can use. Required when |
|
No |
Unset |
Number of requests allowed to wait for a free GPU. Beyond this number, a request is rejected on arrival with |
|
No |
|
How long a request waits for a free GPU before it is rejected with |
|
No |
|
Value sent in the |
|
No |
Derived |
How long a request waits internally for a GPU. Leave this variable unset. The value is derived so that it is longer than |
For additional Boltz-2 tuning variables such as NIM_MODEL_PROFILE, refer to Optimization.
Admission Control#
By default, the NIM accepts every request, and each request waits its turn for a GPU. In a saturated deployment that queue is unbounded, and a client cannot distinguish a slow prediction from one that never starts.
Setting NIM_ENABLE_QUEUE_METRICS to true bounds the queue. Requests beyond NIM_MAX_CONCURRENT_REQUESTS wait, and are rejected with 503 and a Retry-After header after NIM_QUEUE_TIMEOUT_SECONDS elapses, or when NIM_MAX_QUEUE_DEPTH is exceeded. While admission control is enabled, /v1/metrics publishes nim_request_queue_depth, nim_requests_running, nim_request_queue_time_seconds, and nim_requests_shed_total.
docker run -it --rm --name=boltz2 \
--runtime=nvidia \
--gpus 4 \
-e NGC_API_KEY \
-e NIM_ENABLE_QUEUE_METRICS=true \
-e NIM_MAX_CONCURRENT_REQUESTS=4 \
-e NIM_MAX_QUEUE_DEPTH=8 \
-e NIM_QUEUE_TIMEOUT_SECONDS=600 \
-e NIM_QUEUE_RETRY_AFTER_SECONDS=180 \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-p 8000:8000 \
nvcr.io/nim/mit/boltz2:1.10.0
The following table describes the three requirements that admission control cannot infer for you.
Requirement |
Reason |
|---|---|
|
The NIM serves one request per GPU, and its worker pool enforces that limit regardless of this setting. A larger value moves the queue to a place that the metrics cannot observe, and a smaller value leaves GPUs idle. The container refuses to start when the two values disagree, rather than report a limit that does not hold. |
Keep |
Each worker process enforces the concurrency limit separately, so several workers admit a multiple of the intended number of requests. |
Health and documentation endpoints bypass admission control. |
A saturated NIM continues to answer |
NIM Telemetry#
NIM Telemetry collects minimal host and NIM information to help improve performance, reliability, and compatibility across deployments. Collection is disabled by default on RTX GPUs and can be enabled explicitly when you opt in.
NIM_TELEMETRY_MODE#
Controls telemetry collection mode.
0— Disables telemetry.1— Enables baseline collection (hardware and NIM information).
Default: 0
NIM_TELEMETRY_ENABLE_ON_RTX#
Set to true to allow telemetry collection on RTX GPUs. This setting applies only when NIM_TELEMETRY_MODE is not 0 (for example, when baseline collection is enabled on the deployment).
Default: false
For what is collected, retention, and advanced options, refer to NIM Telemetry Settings and NVIDIA’s Privacy Policy.
Volumes#
The following table describes the paths inside the container into which local paths can be mounted.
Container path |
Required |
Notes |
Docker argument example |
|---|---|---|---|
|
Not required, but if this volume is not mounted, the container will do a fresh download of the model each time it is brought up. |
This is the directory within which models are downloaded inside the container. It is very important that this directory can be accessed from inside the container. This can be achieved by setting the permissions of the local directory to |
|
|
Not required |
Destination for exposed prediction artifacts such as confidence JSONs and full PAE/PDE |
|
Logging Configuration#
The Boltz-2 NIM provides logging environment variables to control verbosity for different components. You can configure these when starting the container to adjust the level of detail in logs for debugging or production use.
Example: Running with Logging Configuration#
The following example shows how to run the NIM with standard logging configuration:
export LOCAL_NIM_CACHE=~/.cache/nim
export NGC_API_KEY=<Your NGC API Key>
docker run --rm --name boltz2 --runtime=nvidia \
--shm-size=16G \
-e NGC_API_KEY \
-e NIM_LOG=INFO \
-e NIM_LOG_LEVEL=INFO \
-v $LOCAL_NIM_CACHE:/opt/nim/.cache \
-p 8000:8000 \
nvcr.io/nim/mit/boltz2:1.10.0
Example: Debug Logging for Troubleshooting#
To enable verbose logging for troubleshooting issues, set the log levels to their most detailed settings:
export LOCAL_NIM_CACHE=~/.cache/nim
export NGC_API_KEY=<Your NGC API Key>
docker run --rm --name boltz2 --runtime=nvidia \
--shm-size=16G \
-e NGC_API_KEY \
-e NIM_LOG=DEBUG \
-e NIM_LOG_LEVEL=DEBUG \
-v $LOCAL_NIM_CACHE:/opt/nim/.cache \
-p 8000:8000 \
nvcr.io/nim/mit/boltz2:1.10.0