Observability#

NIM provides Prometheus metrics indicating request statistics. You can use these metrics to create dashboards with Grafana. By default, these metrics are available at http://localhost:8000/v1/metrics.

You can use the following command to retrieve the metrics:

curl -X 'GET' 'http://localhost:8000/v1/metrics'

The following table describes the available metrics.

Category

Metric Name

Description

GPU

gpu_power_usage_watts

GPU instantaneous power, reported in milliwatts

gpu_power_limit_watts

Maximum GPU power limit, reported in milliwatts

gpu_total_energy_consumption_joules

GPU total energy consumption, in joules

gpu_utilization

GPU utilization rate (0.0–1.0)

gpu_memory_total_bytes

Total GPU memory, in bytes

gpu_memory_used_bytes

Used GPU memory, in bytes

Process

process_virtual_memory_bytes

Virtual memory size in bytes

process_resident_memory_bytes

Resident memory size in bytes

process_start_time_seconds

Start time of the process since Unix epoch in seconds

process_cpu_seconds_total

Total user and system CPU time spent in seconds

process_open_fds

Number of open file descriptors

process_max_fds

Maximum number of open file descriptors

Python

python_gc_objects_collected_total

Objects collected during GC

python_gc_objects_uncollectable_total

Uncollectable objects found during GC

python_gc_collections_total

Number of times this generation was collected

python_info

Python platform information

Request

request_count_total

Request counter declared by the shared serving framework. This NIM does not increment it, so the endpoint reports its type but no samples.

request_latency_seconds

Request latency histogram declared by the shared serving framework, on the same basis as the preceding row.

Note

The two request_* series appear in a scrape as HELP and TYPE declarations with no sample lines. They are part of the shared serving framework rather than of this NIM, which serves inference over gRPC. For per-inference counts and latencies, use the Triton metrics described in the next section.

Note

gpu_power_usage_watts and gpu_power_limit_watts report milliwatts despite their names. Divide both by 1000 before charting them in watts.

Triton Metrics#

NIM also exposes Triton Inference Server metrics that provide detailed information about model inference performance, request handling, and system utilization. By default, these Triton metrics are available at http://localhost:9002/metrics.

You can use the following command to retrieve the Triton metrics:

curl localhost:9002/metrics

Note

When running NIM in a container, ensure that port 9002 is properly forwarded by including the -p 9002:9002 flag in your Docker run command.

Every per-model series carries the label model="3DBodyPoseEstimation", which is the name the NIM infers with. Filter on it to separate this model’s series from Triton’s own aggregates, for example nv_inference_request_success{model="3DBodyPoseEstimation"}.

The following are key Triton metrics:

  • Request metrics: Counts of successful and failed inference requests. The NIM sends one inference request per decoded frame, regardless of how many bodies the frame carries, plus about 120 requests per stream to fill and drain the temporal window of the model. A 300-frame clip therefore advances the request counters by about 420.

  • Inference metrics: Request queue times, compute times, and overall request durations.

  • Model metrics: Model loading times, execution counts, and batch statistics.

  • Memory metrics: GPU and CPU memory usage for inference operations.

  • Cache metrics: Response cache hit and miss rates (when caching is enabled).

For comprehensive documentation on all available Triton metrics, their descriptions, and usage examples, refer to Metrics in the Triton Inference Server guide.