Planner Examples
Examples for custom load predictors and the VirtualConnector for non-Kubernetes scaling environments.
Planner-specific examples for advanced configuration and non-Kubernetes integrations. For DGDR manifests, see DGDR Templates. For the full configuration reference, see the Planner Guide.
Custom Load Predictors
Each YAML block in this section is a standalone PlannerConfig. Save the block
as planner.yaml and pass it to
python -m dynamo.planner --config planner.yaml. To use the same fields in a
DGDR, nest them under spec.features.planner.
Warm-starting with Trace Data
Pre-load predictors with historical request patterns before live traffic:
The parser accepts per-request Mooncake JSONL records:
It also accepts dynamo.request.trace.v1 request_end records. The Planner
groups requests into adjustment intervals and computes request count, average
input sequence length (ISL), and average output sequence length (OSL).
Kalman Filter Tuning
For workloads with rapid changes, tune the Kalman filter:
Prophet for Seasonal Workloads
For workloads with daily/weekly patterns:
Power-Aware Budget Scaling
Keep the Planner’s projected GPU power draw within a configured rack/DGD budget.
Per-GPU caps are DGD-owned: authored on each worker component’s podTemplate
annotation (dynamo.nvidia.com/gpu-power-limit), applied to Pods by the
operator, and enforced by the Power Agent. The Planner only reads them and
combines them with total_gpu_power_limit (in its config) to project a budget
and clamp scale-up — it never patches Pods.
The mounted PlannerConfig enables it:
enable_power_awareness requires environment: "kubernetes" and
mode set to disagg, prefill, or decode (agg is not supported).
The Planner caches each annotated component’s cap, effective main-container GPU
count, and node count at startup. DGD admission rejects changes to those fields;
delete and recreate the DGD to change them. Restart the Planner after changing
total_gpu_power_limit.
You must also enable pods/list RBAC for the Planner’s ServiceAccount at
install time. The Planner reads Pod annotations during startup to verify that
power caps have propagated before caching them. Without the permission the
startup settlement check fails. Pass this flag when installing or upgrading the
platform chart:
See the power-aware-budget/ directory in
Dynamo examples for
the full annotation + config contract and its limitations (the budget is a
projected ceiling over requested caps, not a proven hardware limit). Mixed GPU
generations, dynamic cap retargeting, and DRA-backed GPU allocation are not
supported.
Virtual Connector
For non-Kubernetes environments, use the VirtualConnector to communicate scaling decisions:
See the VirtualConnector integration test for a complete example.
Related Documentation
- Planner Guide — Planner configuration reference
- DGDR Templates — DGDR YAML examples
- Profiler Guide — Profiling workflow