Advanced Usage#
This page describes deployment options, request parameters, and runtime behavior for the 3D Body Pose NIM.
The launch commands on this page read NGC_API_KEY from your shell and pull the image from nvcr.io. Export the key and log in with docker login nvcr.io first, as described in Getting Started.
Model Caching#
On first launch, the container downloads the required models from NGC (about 3.5 GB). Mount a host directory as a cache to avoid re-downloading on subsequent runs.
Create the cache directory on the host before the first docker run:
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
Run the container with the cache directory mounted in the appropriate location:
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-e NGC_API_KEY=$NGC_API_KEY \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
nvcr.io/nim/nvidia/body-pose:latest
docker logs -f body-pose-nim
Note
LOCAL_NIM_CACHE can be any directory you own. If the log reports that the model download target is not writable, the directory most likely belongs to another user, typically root because Docker created it on an earlier run; change the block’s first line to a directory directly under your home, for example export LOCAL_NIM_CACHE=~/nim-cache, and run the block again. The container runs as uid 1000, so the directory needs a+rwX, and the files it downloads belong to uid 1000.
To pin the profile that is downloaded and loaded, set NIM_MODEL_PROFILE. For more information, refer to Model Manifest Profiles.
SSL Enablement#
3D Body Pose NIM provides an SSL mode to ensure secure communication between clients and the server by encrypting data in transit. To enable SSL, you must provide the path to the SSL certificate and key files in the container. The following example shows how to do this:
The container expects a CA certificate, a server certificate signed by it, and the matching server key. mtls mode also needs a client certificate signed by the same CA. To generate a self-signed set for local testing, in an empty directory:
mkdir -p certs && cd certs
# 1. CA
openssl req -x509 -newkey rsa:4096 -nodes -days 365 \
-keyout ssl_ca_key.pem -out ssl_ca_cert.pem \
-subj "/CN=body-pose-test-ca"
# 2. Server key + CSR
openssl req -newkey rsa:4096 -nodes \
-keyout ssl_key_server.pem -out server.csr \
-subj "/CN=localhost"
# 3. Server certificate, signed by the CA. The client verifies the server's
# hostname against the certificate, so the SAN is required, not optional.
printf "subjectAltName=DNS:localhost,IP:127.0.0.1\n" > san.cnf
openssl x509 -req -in server.csr -days 365 \
-CA ssl_ca_cert.pem -CAkey ssl_ca_key.pem -CAcreateserial \
-extfile san.cnf \
-out ssl_cert_server.pem
# 4. Client key + CSR (mtls only -- the server rejects a connection with no
# client certificate)
openssl req -newkey rsa:4096 -nodes \
-keyout ssl_key_client.pem -out client.csr \
-subj "/CN=body-pose-client"
# 5. Client certificate, signed by the same CA
openssl x509 -req -in client.csr -days 365 \
-CA ssl_ca_cert.pem -CAkey ssl_ca_key.pem -CAcreateserial \
-out ssl_cert_client.pem
# openssl writes ssl_key_server.pem at mode 600 (owner-only). The server runs
# inside the container as uid 1000 (user nvs), not as the uid that ran openssl,
# so grant that uid, and only that uid, read access; otherwise the server
# refuses to start:
# ERROR:inference:SSL configuration is invalid; refusing to start.
# nimlib.exceptions.SSLConfigurationError: Error verifying SSL files: [Errno 13] Permission denied
# Nothing to do when you are uid 1000. setfacl comes from the acl package.
[ "$(id -u)" = 1000 ] || setfacl -m u:1000:r ssl_key_server.pem
cd ..
Note
This block is POSIX shell, so it runs under sh/dash as well as bash and zsh. It leaves a few scratch files next to the certificates – san.cnf, the two CSRs (server.csr, client.csr), and the CA serial (ssl_ca_cert.srl).
Important
SSL_CERT must be an absolute path. Docker reads a relative -v source as the name of a volume rather than a directory, so a relative value silently creates an empty named volume, mounts it over /opt/nim/crt, and the server then exits reporting that it cannot find a certificate you just generated.
# Run from the directory where you ran the recipe above, so that $PWD/certs is the directory it created.
export SSL_CERT="$PWD/certs"
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-v "$SSL_CERT:/opt/nim/crt/:ro" \
-e NGC_API_KEY=$NGC_API_KEY \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
-e NIM_SSL_MODE="mtls" \
-e NIM_SSL_CA_CERTS_PATH="/opt/nim/crt/ssl_ca_cert.pem" \
-e NIM_SSL_CERT_PATH="/opt/nim/crt/ssl_cert_server.pem" \
-e NIM_SSL_KEY_PATH="/opt/nim/crt/ssl_key_server.pem" \
nvcr.io/nim/nvidia/body-pose:latest
docker logs -f body-pose-nim
NIM_SSL_MODE can be set to mtls, tls, or disabled. If set to mtls, the container uses mutual TLS authentication. If set to tls, the container uses TLS authentication.
For more information, refer to Environment Variables.
Verify the permissions of the SSL certificate and key files on the host machine. The container runs as uid 1000 and cannot start if the files are not readable by that user, for example a key that openssl wrote with mode 600. Grant read access to uid 1000 only, with setfacl -m u:1000:r ssl_key_server.pem, rather than making the private key world-readable. The server certificate must carry a Subject Alternative Name for the hostname that clients connect to.
Note
NIM_SSL_MODE secures the gRPC port (8001) only. The HTTP endpoints on port 8000 stay plaintext.
To connect the sample Python client to an SSL-enabled server, set --ssl-mode to match NIM_SSL_MODE. Pass --ssl-root-cert for TLS or MTLS, and additionally --ssl-cert and --ssl-key for MTLS. Run the client from its scripts folder, with the virtual environment from Basic Inference active. The launch commands above export SSL_CERT, so the first line below keeps that value in the same shell. In a new shell, first run export SSL_CERT="{absolute path of the certs directory the recipe created}"; otherwise the first line stops with a message naming SSL_CERT:
export SSL_CERT="${SSL_CERT:?set SSL_CERT to the absolute path of the certs directory the recipe created}" && \
python body_pose.py --target 127.0.0.1:8001 --video-input ../assets/sample_video.mp4 --bbox-input ../assets/sample_bbox.txt --output body_pose_output.json \
--ssl-mode MTLS \
--ssl-root-cert "$SSL_CERT/ssl_ca_cert.pem" \
--ssl-cert "$SSL_CERT/ssl_cert_client.pem" \
--ssl-key "$SSL_CERT/ssl_key_client.pem"
A client whose --ssl-mode does not match the server fails to connect with UNAVAILABLE, for example a plaintext client against a tls or mtls server, or --ssl-mode TLS against an mtls server.
Multiple Concurrent Inputs#
To enable multi-stream concurrent inference, set NV_AI4M_MAX_CONCURRENCY_PER_GPU to an integer greater than 1 on the server container.
Because Triton distributes the workload equally across all GPUs, the total number of concurrent inputs supported by the server is the number of GPUs multiplied by NV_AI4M_MAX_CONCURRENCY_PER_GPU.
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-e NGC_API_KEY=$NGC_API_KEY \
-e NV_AI4M_MAX_CONCURRENCY_PER_GPU=2 \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
nvcr.io/nim/nvidia/body-pose:latest
docker logs -f body-pose-nim
A request that arrives when every slot is busy waits up to BODY_POSE_REQUEST_ADMISSION_TIMEOUT_SEC (default 15 seconds) for a slot, then fails with RESOURCE_EXHAUSTED. The wait is silent: the server takes the slot before it reads the first message off the stream, so nothing you sent was consumed.
The status detail names the two numbers that decide the outcome:
Server at capacity: all 2 admission slot(s) in use and none freed within 15s (BODY_POSE_REQUEST_ADMISSION_TIMEOUT_SEC). Retrying without changing anything will fail the same way. Serialise your streams, or raise NV_AI4M_MAX_CONCURRENCY_PER_GPU on the server and restart it.
Retrying immediately against an unchanged server fails identically, so treat this as a capacity decision rather than a transient error: either serialise your streams, or raise NV_AI4M_MAX_CONCURRENCY_PER_GPU and restart the container.
Note
Raising the concurrency lets several streams run at once but does not add throughput: on a saturated GPU, the aggregate frame rate is nearly flat from one stream to eight, so the same total is divided into more, slower streams. Refer to Concurrency. The GPU’s concurrent NVDEC session limit also bounds the number of concurrent inputs; refer to Unsupported Architectures.
Focal Length#
The focal_length field in the BodyPoseConfig proto sets the pinhole focal length, in pixels, that the model uses, where fx equals fy.
If the field is unset or set to
0, the NIM uses the SDK default derived from the image size.Supplying the true focal length of the capture camera improves the metric accuracy of
keypoints_3d.
To set the focal length, include the field in the BodyPoseConfig message sent as the first request in the gRPC stream. Run the following from the client package root (nim-clients/body-pose/), where the relative interfaces path resolves, with the virtual environment from Basic Inference active:
import sys
sys.path.insert(0, "interfaces") # where compile_protos.sh writes the stubs
from nvidia.ai4m.body_pose.v1 import body_pose_pb2
config = body_pose_pb2.BodyPoseConfig(focal_length=1200.0)
With the sample client, use the --focal-length argument. Run it from the client’s scripts folder, with the same virtual environment active:
python body_pose.py --target 127.0.0.1:8001 --video-input ../assets/sample_video.mp4 --bbox-input ../assets/sample_bbox.txt --output body_pose_output.json --focal-length 1200
Contact Correction#
The enable_contact field in the BodyPoseConfig proto turns on static-camera contact correction, which runs inverse kinematics to keep the body in contact with the ground plane. It is intended for a fixed camera with the subject standing or walking on a consistent ground plane, and it costs some throughput.
Contact correction is a server setting, fixed at startup: it is off unless the container is started with BODY_POSE_ENABLE_CONTACT=1. A request whose enable_contact value differs from the server’s is processed with the server’s value, and the response carries a warning in its initial metadata, which the sample client prints as Server warning:. Run one server per setting if you need both.
To start a server with contact correction on:
export LOCAL_NIM_CACHE=~/.cache/nim
mkdir -p "$LOCAL_NIM_CACHE"
chmod a+rwX "$LOCAL_NIM_CACHE" 2>/dev/null || true
docker rm -f body-pose-nim 2>/dev/null
docker run -d --name=body-pose-nim \
--runtime=nvidia \
--gpus all \
--shm-size=8GB \
-v "$LOCAL_NIM_CACHE:/opt/nim/.cache" \
-e NGC_API_KEY=$NGC_API_KEY \
-e BODY_POSE_ENABLE_CONTACT=1 \
-e NIM_HTTP_API_PORT=8000 \
-e NIM_GRPC_API_PORT=8001 \
-p 8000:8000 \
-p 8001:8001 \
-p 9002:9002 \
nvcr.io/nim/nvidia/body-pose:latest
docker logs -f body-pose-nim
With the sample client, request contact correction with the --enable-contact argument, or turn it off with --no-enable-contact; omit both to accept the server’s value. Run it from the client’s scripts folder, with the virtual environment from Basic Inference active:
python body_pose.py --target 127.0.0.1:8001 --video-input ../assets/sample_video.mp4 --bbox-input ../assets/sample_bbox.txt --output body_pose_output.json --enable-contact
Confirming Which Value the Server Runs With#
Contact correction only changes the output when a person is in frame, so a clip without one looks identical either way. Read the value off the server instead of inferring it from the output. The server logs it once, after the warmup inference latches it:
enable_contact startup latch done: loaded=True (from BODY_POSE_ENABLE_CONTACT)
The grep below returns that line. With NIM_ENABLE_OTEL=1, the text is the body of an OpenTelemetry log record, so the grep returns the JSON line "body": "enable_contact startup latch done: ..." instead. loaded=True means contact correction is on for every stream this server serves; loaded=False means it is off. Nothing changes the value afterwards.
docker logs body-pose-nim 2>&1 | grep "enable_contact startup latch done"
Service Information in the Response#
Each response stream opens with a one-shot ServiceInfo message on the service_info field of BodyPoseResponse, before any pose result. It reports feature_name, feature_version, model_info, server_request_id, and client_session_id.
Quote server_request_id when reporting an issue; the log lines the pose pipeline writes for the stream carry it as a [GUID=<id>] prefix (a few decoder-lifecycle lines, such as GstVideoReader opened and End-Of-Stream reached., are written without it). In the following snippet, response_iterator is the response stream returned by BodyPoseServiceStub.EstimateBodyPose():
for response in response_iterator:
if response.HasField("service_info"):
print(response.service_info.server_request_id)
elif response.bodies:
print(response.frame_id, len(response.bodies))
client_session_id echoes the value that the client sent in the client-session-id gRPC metadata header, and is empty when the client sends none. The sample client sets it from --client-session-id.
Advanced Tuning Variables#
These variables tune internal timeouts and buffer sizes. The defaults suit every workload described in Performance Results. Change them only to address a specific problem.
ENV |
Default |
Notes |
|---|---|---|
|
|
OpenTelemetry on or off. Set |
|
|
Where OpenTelemetry log records go: |
|
|
Same choice for spans. |
|
|
Same choice for metrics. |
|
Unset |
The OTLP endpoint that receives OpenTelemetry records when one of the exporters above is set to |
|
Unset |
Absolute cap on concurrently admitted requests. When set, it overrides |
|
|
Largest single gRPC message the server accepts, in bytes (20 MiB by default). Raise it only if a request carrying many tracked boxes is rejected with |
|
|
How long a request waits for an admission slot before the server returns |
|
Server-set defaults |
Extra flags for the Triton server launch command. Setting it replaces the defaults, except that |
|
|
How long the server waits for the model files to resolve before exiting. |
|
|
How long the server waits at startup for Triton and the gRPC backend to become healthy. On expiry the container exits non-zero. |
|
|
How long the startup warmup may take. On expiry the container keeps running with |
|
|
Width, in pixels, of the synthetic frame used for startup warmup. |
|
|
Height, in pixels, of the synthetic frame used for startup warmup. |
|
|
Per-inference deadline for a single Triton call. |
|
|
How long the decoder may run without producing a first frame before the request fails with |
|
|
Maximum number of buffered pose results held between the decode pipeline and the gRPC response stream. Larger values tolerate a slower client at the cost of memory. |
|
|
Maximum number of decoded frames held in flight between the decoder and the pipeline. When the video decodes to fewer frames than the annotation names, a shortfall no larger than this depth is also logged as a warning that the decode may have stopped early; the caller gets |
|
|
Upper bound on the iterations used to drain the model’s temporal window at the end of a stream. |