Configuration Alternatives at Runtime#

This section provides a guide on how to run the OpenFold3 NIM with different settings for

  1. port number for http requests

  2. template processing

  3. serving your own model checkpoint

  4. serving requests on multiple GPUs

  5. behavior under load, including inputs too large for the GPU

  6. reproducibility

  7. request size limits

These instructions assume

  1. You have installed and set up Prerequisite Software (Docker, NGC CLI, NGC registry access).

  2. You have pulled the container as described in Getting Started

Start the OpenFold3 NIM#

To start the NIM:

export NGC_API_KEY=<Your NGC API Key>
export LOCAL_NIM_CACHE=~/.cache/nim

docker run --rm --name openfold3 \
    --runtime=nvidia \
    --gpus 'device=0' \
    -e NGC_API_KEY \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 8000:8000 \
    --shm-size=16g \
    nvcr.io/nim/openfold/openfold3:latest

Note

The meaning of the following docker run options are given below

  • The -p option sets the port for the NIM.

  • The -e options define the environment variables, which are passed into the NIM’s container at runtime.

  • The --rm option removes the container and associated file system usage when it exits / terminates.

Using an Alternative Port for OpenFold3 NIM Requests#

If you have other HTTP servers running (for example, other NIMs), you may need to make the 8000 port available by using another port for your NIM. To use an alternative port:

  1. Change the exposed port by setting the -p option.

  2. Set the NIM_HTTP_API_PORT environment variable to the new port.

The following is an example of setting the NIM to run on port 6626:

export NGC_API_KEY=<Your NGC API Key>
export LOCAL_NIM_CACHE=/mount/largedisk/nim/.cache

docker run --rm --name openfold3 \
    --runtime=nvidia \
    --gpus 'device=0' \
    -e NGC_API_KEY \
    -e NIM_HTTP_API_PORT=6626 \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 6626:6626 \
    --shm-size=16g \
    nvcr.io/nim/openfold/openfold3:latest
  • NIM_HTTP_API_PORT=6626 sets the in-container HTTP port.

  • -p 6626:6626 publishes the same port on the host.

Using Your Own Model Checkpoint#

The NIM ships with the OpenFold3 checkpoint it was released with, and downloads it into the NIM cache on first run. You can serve a different checkpoint, typically weights you have fine-tuned yourself, by mounting it into the container and naming it with two environment variables:

Variable

Default

Purpose

MODEL_PATH

the NIM cache

Directory holding the checkpoint.

MODEL_FILE_NAME

of3-p2-155k.pt

Name of the checkpoint file in that directory.

MODEL_FILE_NAME exists so you do not have to rename your file: point the two variables at whatever you produced.

export NGC_API_KEY=<Your NGC API Key>
export LOCAL_NIM_CACHE=~/.cache/nim
export MY_CHECKPOINT_DIR=/data/openfold3-finetuned

docker run --rm --name openfold3 \
    --runtime=nvidia \
    --gpus 'device=0' \
    -e NGC_API_KEY \
    -e MODEL_PATH=/checkpoints \
    -e MODEL_FILE_NAME=my-finetuned.pt \
    -v $MY_CHECKPOINT_DIR:/checkpoints:ro \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 8000:8000 \
    --shm-size=16g \
    nvcr.io/nim/openfold/openfold3:latest

Confirm which checkpoint was loaded from the startup log:

api.di - INFO - Model path: /checkpoints/my-finetuned.pt

What Is and Is Not Checked#

A checkpoint you name must exist. The NIM fails to start and says so if MODEL_PATH and MODEL_FILE_NAME do not resolve to a file, whether from a typo, a volume that did not mount, or a filename that does not match:

FileNotFoundError: No checkpoint at /checkpoints/my-finetuned.pt. MODEL_PATH=...
and MODEL_FILE_NAME=... name a checkpoint explicitly, so the NIM will not fall
back to the one in its cache — that would serve different weights than you
asked for.

This is deliberate. Quietly serving the shipped checkpoint instead would look identical from the outside: the server starts, predictions return, confidence scores look ordinary. A fine-tune-and-re-serve workflow would then measure the model it was trying to replace. Leave both variables unset to use the shipped checkpoint.

The architecture is not checked. The NIM loads the checkpoint into the OpenFold3 model it was built with, so the checkpoint must come from the same architecture, meaning fine-tuned weights rather than a different model. A mismatch surfaces as a PyTorch state-dict error at startup rather than a message from the NIM.

No conversion is performed. Supply a PyTorch checkpoint in the same format as the one the NIM ships. The NIM does not train or fine-tune; produce the weights with your own training code and serve them here.

Note

Accuracy figures published for the OpenFold3 NIM describe the shipped checkpoint. Predictions from your own weights are yours to validate.

Template Configuration Options#

Configuring Template Chain Selection Threshold#

When using structural templates, the NIM validates that template chains have sufficient similarity to your query sequence. You can control the minimum score threshold using the CIF_DIRECT_MIN_SCORE environment variable.

Score Calculation:

score = sequence_identity × query_coverage

The variables are:

  • sequence_identity: Fraction of aligned residues that match between template and query

  • query_coverage: Fraction of the query sequence covered by the alignment

Default Value: 0.1 (10% identity × coverage)

How It Works:

  • With chain_id specified: The specified chain in the input CIF string is validated against the threshold, and the template is rejected if its score is below the threshold.

  • Without chain_id (automatic selection): All chains in the input CIF string are evaluated, and the best-matching chain is selected only if its score meets the threshold; otherwise, no chain from the input CIF string is used.

Note

The chain_id parameter is specified in the inference request payload when providing structural templates. For example, in the templates field of your API request, you can include "chain_id": "A" to specify which chain from the input CIF string should be used for alignment validation.

Configuration Example:

export NGC_API_KEY=<Your NGC API Key>
export LOCAL_NIM_CACHE=~/.cache/nim

docker run --rm --name openfold3 \
    --runtime=nvidia \
    --gpus 'device=0' \
    -e NGC_API_KEY \
    -e CIF_DIRECT_MIN_SCORE=0.2 \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 8000:8000 \
    --shm-size=16g \
    nvcr.io/nim/openfold/openfold3:latest
  • CIF_DIRECT_MIN_SCORE=0.2 sets the minimum score threshold to 0.2.

Guidance:

  • Lower values (0.05-0.15): More permissive, includes template sequence with very few residues in common with the query sequence.

  • Default value (0.1): Balanced approach for most use cases.

  • Higher values (0.2-0.4): More restrictive, includes only template sequences with many residues in common with the query sequence.

Serving Requests on Multiple GPUs#

Starting in 1.6.0, the NIM detects every GPU made visible to the container and starts one inference process per GPU. Each process loads its own copy of the model and serves one request at a time; requests are handed to whichever process is free. A prediction never spans more than one GPU, so this increases throughput rather than reducing the latency of any single request.

To use it, give the container more than one GPU:

export NGC_API_KEY=<Your NGC API Key>
export LOCAL_NIM_CACHE=~/.cache/nim

docker run --rm --name openfold3 \
    --runtime=nvidia \
    --gpus all \
    -e NGC_API_KEY \
    -v $LOCAL_NIM_CACHE:/opt/nim/.cache \
    -p 8000:8000 \
    --shm-size=16g \
    nvcr.io/nim/openfold/openfold3:latest

Note

Each process loads the model independently, so startup time and host memory both scale with the number of GPUs. The container reports ready only once every process has finished loading.

Behavior Under Load#

Requests beyond the number of GPUs are queued. Once the queue is also full, the NIM rejects further requests with HTTP 503:

{"message": "All GPU workers are busy and the queue is full; retry shortly."}

The status is 503 rather than 429 because the NIM is not rate-limiting a particular caller. It has no free worker and a bounded queue. Clients should retry after a short delay, and load balancers should route elsewhere.

Abandoned Requests#

A client that gives up on a request, through a timeout or a closed connection, leaves the NIM holding work nobody reads. A queued request whose caller has disconnected is discarded when its turn comes rather than run, and the response is recorded as HTTP 499. Nothing receives that status because the connection is already closed. It exists so that the outcome is attributed to the client rather than counted as a server error.

This matters most when one long prediction is in front of a queue: without it, every request behind an abandoned one waits for a result that is thrown away.

Note

A prediction that has already started runs to completion even if its caller disappears. Set a client timeout generous enough for the sequence lengths you send, referring to Table 1 for measured runtimes, rather than relying on a short one to free the GPU.

Inputs Too Large for the GPU#

A prediction that exhausts GPU memory is reported as HTTP 413, with a message naming what the request asked for and what the allocator wanted:

{"message": "Out of memory on NVIDIA H100 80GB HBM3 while predicting this request (4096 residues across 1 molecule, 1 diffusion sample). Send a shorter input, ask for fewer diffusion samples, drop structural templates, or use a GPU with more memory. Allocator detail: CUDA out of memory. Tried to allocate 16.00 GiB ..."}

The status is 413 rather than 503 or 500 because retrying will not help. The same input on the same GPU exhausts it every time. What can change is the request or the hardware.

There is no fixed maximum sequence length to check against, which is why this is reported when it happens rather than rejected up front. Memory use depends on the number of residues, whether structural templates are supplied, and how many diffusion samples are requested. A length that succeeds without templates can exhaust the same GPU with them. Refer to Table 1 for input sizes measured on each SKU.

Reproducibility#

By default the NIM seeds its sampler per request, so the same input returns the same structures on every call. Requesting several diffusion samples still gives several different structures within that one response; it is repeat calls that agree.

Variable

Default

Purpose

NIM_INFERENCE_SEED

42

Seed applied at the start of each request.

Set NIM_INFERENCE_SEED= (empty) to restore unseeded sampling, where repeat calls vary. Set it to any other integer to sample a different fixed trajectory.

Note

Structures are reproducible for a given NIM version on a given GPU. Different GPU architectures do not produce bit-identical results, and neither do different releases, because floating-point kernels differ. The seed removes run-to-run variation, not hardware variation.

Configuration#

Variable

Default

Purpose

NIM_PROCESS_POOL

1

Set to 0 to disable the pool and serve from a single process.

NIM_GPU_IDS

all visible

Comma-separated GPU indices to use, for example 0,1.

NIM_MAX_QUEUED_REQUESTS

10

Requests allowed to wait before the NIM returns 503.

NIM_WORKER_STARTUP_TIMEOUT

1200

Seconds to wait for a worker process to load the model.

Response Snapshots#

The NIM can write a copy of each completed prediction to OUTPUT_DIR as a JSON file. This is a debugging aid and is off by default. The prediction is returned in the HTTP response, and nothing in the NIM reads the snapshots back.

Variable

Default

Purpose

NIM_DUMP_RESPONSES

0

Set to 1 to write a JSON snapshot of every response to OUTPUT_DIR.

OUTPUT_DIR

/opt/nim/output

Where snapshots are written when enabled.

Filenames are generated by the server and are not predictable: response_<id>_<random>.json, where <id> is the request’s request_id reduced to letters, digits, ., - and _. The random suffix means two requests never overwrite one another, including two requests that send the same request_id. Locate a snapshot by globbing the prefix rather than by constructing the name.

Writing a snapshot never affects the response. Each file is written to a temporary name and renamed into place, so a reader sees either nothing or a complete document; if the directory is missing, read-only or full, the NIM logs a warning and still returns the prediction.

Request Size Limits#

The NIM validates the size of an incoming request before running it. These limits guard against inputs large enough to exhaust GPU memory; raise them only if you have verified the hardware can handle the workload.

Variable

Default

Purpose

NIM_MAX_POLYMER_INPUTS

12

Maximum number of polymer chains (protein, DNA, RNA) per request.

NIM_MAX_LIGAND_INPUTS

20

Maximum number of ligands per request.

NIM_MAX_POLYMER_LENGTH

102400

Maximum length of any single polymer chain, in residues.

Note

The NIM does not impose a maximum total sequence length. What a given GPU can actually complete depends on its memory; refer to Table 1 for the input sizes measured on each SKU.