Getting Started#
This page shows how to authenticate with NGC, launch the 3D Body Pose NIM container, and configure its runtime.
Prerequisites#
To ensure that you have the supported hardware and software stack, check the Support Matrix.
NGC Authentication#
An NGC API key is required to pull the container image and download models. Complete the following steps to create the key, export it, and log in to the NVIDIA Container Registry.
Generate an API Key#
An NGC API key is required to access NGC resources. You can generate a key at https://org.ngc.nvidia.com/setup/api-keys.
When creating an NGC API Personal key, ensure that at least NGC Catalog is selected from the Services Included dropdown. You can include more services if this key is to be reused for other purposes.
Note
Personal keys allow you to configure an expiration date, revoke or delete the key using an action button, and rotate the key as needed. For more information about key types, refer to NGC API Keys in the NGC User Guide.
Export the NGC API Key#
Pass the value of the API key to the docker run command in the next section as the NGC_API_KEY environment variable to download the appropriate models and resources when starting the NIM.
The simplest way to create the NGC_API_KEY environment variable is to export it in your terminal:
export NGC_API_KEY=<value>
To make the key available in later sessions, store it in a file that only you can read and export it from that file in your shell startup file:
mkdir -p ~/.ngc && (umask 077; printf '%s\n' '<value>' > ~/.ngc/api-key)
export NGC_API_KEY="$(cat ~/.ngc/api-key)" # add this line to ~/.bashrc or ~/.zshrc
Note
A password manager works the same way, for example export NGC_API_KEY="$(pass show nvidia/ngc)".
Docker Login to NGC#
To pull the NIM container image from NGC, first authenticate with the NVIDIA Container Registry with the following command:
echo "$NGC_API_KEY" | docker login nvcr.io --username '$oauthtoken' --password-stdin
Use $oauthtoken as the username and NGC_API_KEY as the password. The $oauthtoken username is a special name that indicates that you will authenticate with an API key and not a user name and password.
Launching the NIM Container#
The following commands launch the 3D Body Pose NIM container with the gRPC service. For a list of parameters, refer to Runtime Parameters for the Container.
Create the host cache directory yourself before the first launch so the model download persists across runs. If Docker creates the directory for you, it is owned by root and the container exits with a permission error. LOCAL_NIM_CACHE can be any directory you own.
The container runs as uid 1000, so the directory needs a+rwX. If the container exits reporting that the model download target is not writable, the directory most likely belongs to another user, so change the first line of the launch commands to a directory directly under your home, for example export LOCAL_NIM_CACHE=~/nim-cache, and run them again. For details, refer to Model Caching.
Launch the container:
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-e NGC_API_KEY=$NGC_API_KEY \
-e NV_AI4M_MAX_CONCURRENCY_PER_GPU=1 \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
nvcr.io/nim/nvidia/body-pose:latest
docker logs -f body-pose-nim
Note
The flag --gpus all is used to assign all available GPUs to the NIM container.
To assign specific GPUs to the NIM container (in case of multiple GPUs available in your machine), use --gpus '"device=0,1,2..."'.
The first launch takes longer than later ones: the container image is several gigabytes to pull, and on the first start the server downloads the model into the mounted cache before it binds its ports — the startup log reads holds no engines yet while it does. Later launches from a populated cache reach SERVING much faster.
If the NIM launch is successful, you get a response similar to the following:
INFO: Started server process [58]
INFO: Waiting for application startup.
INFO: Application startup complete.
INFO: Uvicorn running on http://0.0.0.0:8000 (Press CTRL+C to quit)
Triton server is ready
The inference server’s own log lines follow. The line Health: NOT_SERVING appears early, when the gRPC service starts and before the warmup. The final startup lines include, in this order, Triton warmup done in {seconds}s (engines loaded), enable_contact startup latch done: loaded={bool} (from BODY_POSE_ENABLE_CONTACT), and Health: SERVING. These lines have this shape:
[INFO AI4M BASE LOGGER ... PID:{pid}] Health: SERVING
The NIM launcher’s own lines, such as Waiting for backend readiness, use the plain form INFO:inference:{message} instead.
OpenTelemetry is off by default. To write these lines as OpenTelemetry log records instead, one JSON object per message among span and metric records, add -e NIM_ENABLE_OTEL=1 to the docker run command. See Advanced Tuning Variables for the OpenTelemetry variables.
The NIM runs a warmup pass before it reports ready. Until the warmup completes, /v1/health/ready returns 503 and the gRPC port does not accept connections yet; once it does, the gRPC health service reports SERVING. Poll /v1/health/ready, or retry the gRPC health check until it connects, before sending inference requests. Then run the Quick Test before the full sample.
Model Manifest Profiles#
By default, the NIM selects the model profile that matches the compute capability of the detected GPU. One profile is published per architecture.
GPU Architecture (compute capability) |
GPUs |
|---|---|
Ampere (cc 8.0) |
A100, A30 |
Ampere (cc 8.6) |
A10G, A40, RTX 3090 |
Ada (cc 8.9) |
L4, L40S, RTX 6000 Ada Generation, RTX 4090 |
Hopper (cc 9.0) |
H100, H200 |
Blackwell (cc 10.0) |
B200 |
Blackwell (cc 12.0) |
RTX 5090, NVIDIA RTX PRO 6000 Blackwell Server Edition |
NIM_MODEL_PROFILE is optional. To pin a profile, first list the profiles in the image you pulled with the list-model-profiles utility:
docker run --rm --runtime=nvidia --gpus all \
-e NGC_API_KEY=$NGC_API_KEY \
nvcr.io/nim/nvidia/body-pose:latest list-model-profiles
Then add the ID to the launch command:
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
export MODEL_PROFILE_ID=<enter_model_profile_id>
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-e NGC_API_KEY=$NGC_API_KEY \
-e NV_AI4M_MAX_CONCURRENCY_PER_GPU=1 \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-e NIM_MODEL_PROFILE=$MODEL_PROFILE_ID \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
nvcr.io/nim/nvidia/body-pose:latest
docker logs -f body-pose-nim
Note
If NIM_MODEL_PROFILE is set, ensure that the GPU architecture it targets matches the hardware. With an incorrect profile ID the container exits after NV_AI4M_MODEL_READY_TIMEOUT_S (default 120 seconds) with Model validation failed after <n>s (gpu_cc=<cc>, ...). If left unset, the NIM selects a matching profile automatically.
Environment Variables#
The following table describes the environment variables that can be passed into a NIM as a -e argument added to a docker run command:
ENV |
Required? |
Default |
Notes |
|---|---|---|---|
|
Yes |
None |
You must set this variable to the value of your personal NGC API key. |
|
Optional |
|
Location (in container) where the container caches model artifacts. |
|
No |
|
Publish the NIM gRPC service to the prescribed port inside the container. Adjust the port passed to the |
|
No |
|
Publish the NIM HTTP service to the prescribed port inside the container. Supported endpoints are |
|
Optional |
None |
Pin the model profile to download and load for your GPU. If unset, the NIM selects a profile automatically. For more about |
|
No |
|
Set SSL security on the gRPC endpoint to |
|
No |
|
Set the path to CA root certificate inside the NIM. This is required only when |
|
No |
|
Set the path to the server’s public SSL certificate inside the NIM. This is required only when an SSL mode is enabled. |
|
No |
|
Set the path to the server’s private key inside the NIM. This is required only when an SSL mode is enabled. |
|
No |
|
Number of concurrent inference requests to be supported by the NIM server per GPU. Total concurrency = |
|
No |
|
Maximum size, in bytes, of an input video accepted by the NIM. The default is 1 GiB. |
|
No |
|
Log verbosity of the NIM service. One of |
|
No |
|
Start the server with contact correction enabled. Set to |
For the remaining tuning variables, refer to Advanced Tuning Variables.
Runtime Parameters for the Container#
The following table describes the docker run flags used to launch the 3D Body Pose NIM container.
Flags |
Description |
|---|---|
|
Run the container in the background. Use |
|
Delete the container after it stops (refer to docker container run). |
|
Give a name to the NIM container. Use any preferred value. |
|
Ensure NVIDIA drivers are accessible in the container. |
|
Expose NVIDIA GPUs inside the container. On a host with multiple GPUs, you can expose specific GPUs instead. Refer to GPU Enumeration. |
|
Allocate host memory for multi-process communication. |
|
Mount a host directory as the model cache so that the model download persists across runs. |
|
Provide the container with the token necessary to download adequate models and resources from NGC. Refer to NGC Authentication. |
|
Controls the number of concurrent inference requests processed simultaneously per GPU. Default: 1. Higher values enable parallel request processing but can reduce individual request performance due to resource sharing. |
|
Ports published by the container are directly accessible on the host port. |
Stopping the Container#
The launch commands on this page do not use --rm, so stop the container and then
remove it:
docker stop body-pose-nim
docker rm body-pose-nim