Getting Started#

This page shows how to authenticate with NGC, launch the 3D Body Pose NIM container, and configure its runtime.

Prerequisites#

To ensure that you have the supported hardware and software stack, check the Support Matrix.

NGC Authentication#

An NGC API key is required to pull the container image and download models. Complete the following steps to create the key, export it, and log in to the NVIDIA Container Registry.

Generate an API Key#

An NGC API key is required to access NGC resources. You can generate a key at https://org.ngc.nvidia.com/setup/api-keys.

When creating an NGC API Personal key, ensure that at least NGC Catalog is selected from the Services Included dropdown. You can include more services if this key is to be reused for other purposes.

Note

Personal keys allow you to configure an expiration date, revoke or delete the key using an action button, and rotate the key as needed. For more information about key types, refer to NGC API Keys in the NGC User Guide.

Export the NGC API Key#

Pass the value of the API key to the docker run command in the next section as the NGC_API_KEY environment variable to download the appropriate models and resources when starting the NIM.

The simplest way to create the NGC_API_KEY environment variable is to export it in your terminal:

export NGC_API_KEY=<value>

To make the key available in later sessions, store it in a file that only you can read and export it from that file in your shell startup file:

mkdir -p ~/.ngc && (umask 077; printf '%s\n' '<value>' > ~/.ngc/api-key)
export NGC_API_KEY="$(cat ~/.ngc/api-key)"   # add this line to ~/.bashrc or ~/.zshrc

Note

A password manager works the same way, for example export NGC_API_KEY="$(pass show nvidia/ngc)".

Docker Login to NGC#

To pull the NIM container image from NGC, first authenticate with the NVIDIA Container Registry with the following command:

echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin

Use $oauthtoken as the username and NGC_API_KEY as the password. The $oauthtoken username is a special name that indicates that you will authenticate with an API key and not a user name and password.

Launching the NIM Container#

The following commands launch the 3D Body Pose NIM container with the gRPC service. For a list of parameters, refer to Runtime Parameters for the Container.

Create the host cache directory yourself before the first launch so the model download persists across runs. If Docker creates the directory for you, it is owned by root and the container exits with a permission error. LOCAL_NIM_CACHE can be any directory you own.

The container runs as uid 1000, so the directory needs a+rwX. If the container exits reporting that the model download target is not writable, the directory most likely belongs to another user, so change the first line of the launch commands to a directory directly under your home, for example export LOCAL_NIM_CACHE=~/nim-cache, and run them again. For details, refer to Model Caching.

Launch the container:

export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NV_AI4M_MAX_CONCURRENCY_PER_GPU=1 \
  -e NIM_HTTP_API_PORT=8000 \
  -e NIM_GRPC_API_PORT=8001 \
  -p 8000:8000 \
  -p 8001:8001 \
  -p 9002:9002 \
  nvcr.io/nim/nvidia/body-pose:latest

docker logs -f body-pose-nim

Note

The flag --gpus all is used to assign all available GPUs to the NIM container. To assign specific GPUs to the NIM container (in case of multiple GPUs available in your machine), use --gpus '"device=0,1,2..."'.

The first launch takes longer than later ones: the container image is several gigabytes to pull, and on the first start the server downloads the model into the mounted cache before it binds its ports — the startup log reads holds no engines yet while it does. Later launches from a populated cache reach SERVING much faster.

If the NIM launch is successful, you get a response similar to the following:

INFO:     Started server process [58]
INFO:     Waiting for application startup.
INFO:     Application startup complete.
INFO:     Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
Triton server is ready

The inference server’s own log lines follow. The line Health: NOT_SERVING appears early, when the gRPC service starts and before the warmup. The final startup lines include, in this order, Triton warmup done in {seconds}s (engines loaded), enable_contact startup latch done: loaded={bool} (from BODY_POSE_ENABLE_CONTACT), and Health: SERVING. These lines have this shape:

[INFO AI4M BASE LOGGER ... PID:{pid}] Health: SERVING

The NIM launcher’s own lines, such as Waiting for backend readiness, use the plain form INFO:inference:{message} instead.

OpenTelemetry is off by default. To write these lines as OpenTelemetry log records instead, one JSON object per message among span and metric records, add -e NIM_ENABLE_OTEL=1 to the docker run command. See Advanced Tuning Variables for the OpenTelemetry variables.

The NIM runs a warmup pass before it reports ready. Until the warmup completes, /v1/health/ready returns 503 and the gRPC port does not accept connections yet; once it does, the gRPC health service reports SERVING. Poll /v1/health/ready, or retry the gRPC health check until it connects, before sending inference requests. Then run the Quick Test before the full sample.

Model Manifest Profiles#

By default, the NIM selects the model profile that matches the compute capability of the detected GPU. One profile is published per architecture.

GPU Architecture (compute capability)

GPUs

Ampere (cc 8.0)

A100, A30

Ampere (cc 8.6)

A10G, A40, RTX 3090

Ada (cc 8.9)

L4, L40S, RTX 6000 Ada Generation, RTX 4090

Hopper (cc 9.0)

H100, H200

Blackwell (cc 10.0)

B200

Blackwell (cc 12.0)

RTX 5090, NVIDIA RTX PRO 6000 Blackwell Server Edition

NIM_MODEL_PROFILE is optional. To pin a profile, first list the profiles in the image you pulled with the list-model-profiles utility:

docker run --rm --runtime=nvidia --gpus all \
  -e NGC_API_KEY=$NGC_API_KEY \
  nvcr.io/nim/nvidia/body-pose:latest list-model-profiles

Then add the ID to the launch command:

export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
export MODEL_PROFILE_ID=<enter_model_profile_id>

docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
  --runtime=nvidia \
  --gpus all \
  --shm-size=8GB \
  -v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
  -e NGC_API_KEY=$NGC_API_KEY \
  -e NV_AI4M_MAX_CONCURRENCY_PER_GPU=1 \
  -e NIM_HTTP_API_PORT=8000 \
  -e NIM_GRPC_API_PORT=8001 \
  -e NIM_MODEL_PROFILE=$MODEL_PROFILE_ID \
  -p 8000:8000 \
  -p 8001:8001 \
  -p 9002:9002 \
  nvcr.io/nim/nvidia/body-pose:latest

docker logs -f body-pose-nim

Note

If NIM_MODEL_PROFILE is set, ensure that the GPU architecture it targets matches the hardware. With an incorrect profile ID the container exits after NV_AI4M_MODEL_READY_TIMEOUT_S (default 120 seconds) with Model validation failed after <n>s (gpu_cc=<cc>, ...). If left unset, the NIM selects a matching profile automatically.

Environment Variables#

The following table describes the environment variables that can be passed into a NIM as a -e argument added to a docker run command:

ENV

Required?

Default

Notes

NGC_API_KEY

Yes

None

You must set this variable to the value of your personal NGC API key.

NIM_CACHE_PATH

Optional

/opt/nim/.cache

Location (in container) where the container caches model artifacts.

NIM_GRPC_API_PORT

No

8001

Publish the NIM gRPC service to the prescribed port inside the container. Adjust the port passed to the -p/--publish flag of docker run accordingly (for example, -p $NIM_GRPC_API_PORT:$NIM_GRPC_API_PORT). The left side of the colon (:) is the port of your host address and does not need to match $NIM_GRPC_API_PORT. The right side of the colon (:) is the port inside the container, which must match NIM_GRPC_API_PORT (or 8001 if not set).

NIM_HTTP_API_PORT

No

8000

Publish the NIM HTTP service to the prescribed port inside the container. Supported endpoints are /v1/health/ready and /v1/health/live (readiness and liveness probes), /v1/license (returns the license information), /v1/metadata (returns metadata including asset information, license information, model information, and version), /v1/models (returns the loaded models), and /v1/metrics (exposes Prometheus metrics through an ASGI app endpoint).

NIM_MODEL_PROFILE

Optional

None

Pin the model profile to download and load for your GPU. If unset, the NIM selects a profile automatically. For more about NIM_MODEL_PROFILE, refer to Model Manifest Profiles.

NIM_SSL_MODE

No

disabled

Set SSL security on the gRPC endpoint to tls or mtls. Defaults to unsecured endpoint.

NIM_SSL_CA_CERTS_PATH

No

/opt/nim/crt/ssl_ca_cert.pem

Set the path to CA root certificate inside the NIM. This is required only when NIM_SSL_MODE is mtls.

NIM_SSL_CERT_PATH

No

/opt/nim/crt/ssl_cert_server.pem

Set the path to the server’s public SSL certificate inside the NIM. This is required only when an SSL mode is enabled.

NIM_SSL_KEY_PATH

No

/opt/nim/crt/ssl_key_server.pem

Set the path to the server’s private key inside the NIM. This is required only when an SSL mode is enabled.

NV_AI4M_MAX_CONCURRENCY_PER_GPU

No

1

Number of concurrent inference requests to be supported by the NIM server per GPU. Total concurrency = NV_AI4M_MAX_CONCURRENCY_PER_GPU × number of GPUs.

NV_AI4M_MAX_INPUT_FILE_SIZE_BYTES

No

1073741824

Maximum size, in bytes, of an input video accepted by the NIM. The default is 1 GiB.

NV_AI4M_LOG_LEVEL

No

INFO

Log verbosity of the NIM service. One of DEBUG, INFO, WARNING, ERROR, CRITICAL.

BODY_POSE_ENABLE_CONTACT

No

0

Start the server with contact correction enabled. Set to 1 to enable. Refer to Contact Correction.

For the remaining tuning variables, refer to Advanced Tuning Variables.

Runtime Parameters for the Container#

The following table describes the docker run flags used to launch the 3D Body Pose NIM container.

Flags

Description

-d

Run the container in the background. Use docker logs -f body-pose-nim to follow its output.

--rm

Delete the container after it stops (refer to docker container run).

--name=container-name

Give a name to the NIM container. Use any preferred value.

--runtime=nvidia

Ensure NVIDIA drivers are accessible in the container.

--gpus all

Expose NVIDIA GPUs inside the container. On a host with multiple GPUs, you can expose specific GPUs instead. Refer to GPU Enumeration.

--shm-size=8GB

Allocate host memory for multi-process communication.

-v $LOCAL_NIM_CACHE:/opt/nim/.cache

Mount a host directory as the model cache so that the model download persists across runs.

-e NGC_API_KEY=$NGC_API_KEY

Provide the container with the token necessary to download adequate models and resources from NGC. Refer to NGC Authentication.

-e NV_AI4M_MAX_CONCURRENCY_PER_GPU

Controls the number of concurrent inference requests processed simultaneously per GPU. Default: 1. Higher values enable parallel request processing but can reduce individual request performance due to resource sharing.

-p <host_port>:<container_port>

Ports published by the container are directly accessible on the host port.

Stopping the Container#

The launch commands on this page do not use --rm, so stop the container and then remove it:

docker stop body-pose-nim
docker rm body-pose-nim