Observability#
NIM provides Prometheus metrics indicating request statistics. You can use these metrics to create dashboards with Grafana. By default, these metrics are available at http://localhost:8000/v1/metrics.
You can use the following command to retrieve the metrics:
curl -X 'GET' 'http://localhost:8000/v1/metrics'
The following table describes the available metrics.
Category |
Metric Name |
Description |
|---|---|---|
GPU |
gpu_power_usage_watts |
GPU instantaneous power, reported in milliwatts |
gpu_power_limit_watts |
Maximum GPU power limit, reported in milliwatts |
|
gpu_total_energy_consumption_joules |
GPU total energy consumption, in joules |
|
gpu_utilization |
GPU utilization rate (0.0–1.0) |
|
gpu_memory_total_bytes |
Total GPU memory, in bytes |
|
gpu_memory_used_bytes |
Used GPU memory, in bytes |
|
Process |
process_virtual_memory_bytes |
Virtual memory size in bytes |
process_resident_memory_bytes |
Resident memory size in bytes |
|
process_start_time_seconds |
Start time of the process since Unix epoch in seconds |
|
process_cpu_seconds_total |
Total user and system CPU time spent in seconds |
|
process_open_fds |
Number of open file descriptors |
|
process_max_fds |
Maximum number of open file descriptors |
|
Python |
python_gc_objects_collected_total |
Objects collected during GC |
python_gc_objects_uncollectable_total |
Uncollectable objects found during GC |
|
python_gc_collections_total |
Number of times this generation was collected |
|
python_info |
Python platform information |
|
Request |
request_count_total |
Request counter declared by the shared serving framework. This NIM does not increment it, so the endpoint reports its type but no samples. |
request_latency_seconds |
Request latency histogram declared by the shared serving framework, on the same basis as the preceding row. |
Note
The two request_* series appear in a scrape as HELP and TYPE declarations with no sample lines. They are part of the shared serving framework rather than of this NIM, which serves inference over gRPC. For per-inference counts and latencies, use the Triton metrics described in the next section.
Note
gpu_power_usage_watts and gpu_power_limit_watts report milliwatts despite their names. Divide both by 1000 before charting them in watts.
Triton Metrics#
NIM also exposes Triton Inference Server metrics that provide detailed information about model inference performance, request handling, and system utilization. By default, these Triton metrics are available at http://localhost:9002/metrics.
You can use the following command to retrieve the Triton metrics:
curl localhost:9002/metrics
Note
When running NIM in a container, ensure that port 9002 is properly forwarded by including the -p 9002:9002 flag in your Docker run command.
Every per-model series carries the label model="3DBodyPoseEstimation", which is the name the NIM infers with. Filter on it to separate this model’s series from Triton’s own aggregates, for example nv_inference_request_success{model="3DBodyPoseEstimation"}.
The following are key Triton metrics:
Request metrics: Counts of successful and failed inference requests. The NIM sends one inference request per decoded frame, regardless of how many bodies the frame carries, plus about 120 requests per stream to fill and drain the temporal window of the model. A 300-frame clip therefore advances the request counters by about 420.
Inference metrics: Request queue times, compute times, and overall request durations.
Model metrics: Model loading times, execution counts, and batch statistics.
Memory metrics: GPU and CPU memory usage for inference operations.
Cache metrics: Response cache hit and miss rates (when caching is enabled).
For comprehensive documentation on all available Triton metrics, their descriptions, and usage examples, refer to Metrics in the Triton Inference Server guide.