Overview#

NVIDIA 3D Body Pose NIM estimates the 3D pose of every tracked person in a video. For each tracked body in each frame, the NIM returns 2D keypoints in image pixels, 3D joint positions in camera coordinates, per-joint confidence scores, joint rotations, a rest pose, and a root pose. For the numbered joint layout, refer to Output Format.

NVIDIA 3D Body Pose NIM models are built on the NVIDIA software platform, incorporating CUDA, TensorRT, and Triton to offer ready-to-use GPU acceleration. A pose is estimated once per tracked box, so inference cost scales with the number of bodies in frame; refer to Throughput by GPU and Body Count.

Note

The NIM does not detect bodies. You supply a tracked bounding box for each person in each frame, from any detector and tracker that assigns a stable tracking ID. For the annotation layout, refer to Tracked Bounding-Box Annotation Format.

Architecture#

The 3D Body Pose service is built on NVIDIA’s software platform:

  • CUDA for GPU-accelerated processing.

  • TensorRT for optimized neural network inference.

  • Triton Inference Server for efficient model serving and scaling.

  • NVIDIA Augmented Reality (AR) SDK backend for 3D body pose estimation.

The system uses an architecture that performs the following actions:

  1. Accepts a compressed video file streamed as byte chunks, together with the tracked bounding-box annotation for its frames. Refer to Input Media for the accepted formats.

  2. Decodes the input video frame by frame using GPU-accelerated GStreamer (NVDEC).

  3. Runs 3D body pose estimation for each tracked box in each frame through the AR SDK backend on Triton.

  4. Streams back pose results, each tagged with the index of the decoded input frame.

Pose results are delayed by the model’s temporal window; refer to Delayed Outputs.

Input Modes#

The 3D Body Pose NIM supports two modes of operation:

  • Streaming Mode (recommended): In this mode, the input file is streamed to the NIM in chunks. As the chunks arrive, the NIM decodes and runs inference incrementally and streams results back to the client, even before the whole input file is received by the NIM. The NIM automatically detects streamable videos and enables this mode. For best performance, use streamable videos as input.

  • Transactional Mode: The entire input file is buffered by the NIM before processing begins. The NIM automatically falls back to this mode when the input video is not streamable.

For detailed information about when to use each mode and how to convert between file formats, refer to the Input Modes section of Basic Inference.

Important

Constant frame rate is required and is enforced. Tracked boxes are addressed by frame_id, so a variable-frame-rate clip – or a constant-rate clip with a dropped frame – breaks the frame-to-timestamp mapping your tracker assumed. The NIM rejects the request with INVALID_ARGUMENT rather than returning poses matched to the wrong boxes. Re-encode before sending: ffmpeg -i in.mp4 -r 30 -c:v libx264 -pix_fmt yuv420p out.mp4 (the -pix_fmt converts a 4:4:4 or high-bit-depth source to the 4:2:0 8-bit this NIM accepts; see Support Matrix).

Try It Out#

Try the NVIDIA 3D Body Pose NIM at build.nvidia.com/nvidia/body-pose.

To experience the NVIDIA 3D Body Pose NIM API without having to host your own servers, use the Try API feature, which uses the NVIDIA Cloud Function backend.