nemo_rl.algorithms.single_controller_utils.utils#
Helpers used by SingleControllerActor.
Module Contents#
Classes#
One advantage-stage call’s reward sums, already reduced. |
|
One advantage-stage call’s token-masked advantage moments. |
Functions#
Reduce per-microbatch metric lists into step-level scalars. |
|
Reduce per-step accumulators from _advantage_stage into step scalars. |
|
Reduce sequence-error metrics across streaming chunks. |
|
Overwrite flagged token advantages while leaving valid tokens unchanged. |
|
Read a tensor column from a TensorDict, depadding if nested. |
|
Drop a trailing dim of size 1 if present. |
|
Pack tensors for DataPlane put, re-nesting jagged rows when needed. |
Data#
API#
- nemo_rl.algorithms.single_controller_utils.utils._MB_METRIC_MIN: frozenset[str]#
‘frozenset(…)’
- nemo_rl.algorithms.single_controller_utils.utils._MB_METRIC_MAX: frozenset[str]#
‘frozenset(…)’
- nemo_rl.algorithms.single_controller_utils.utils._MB_METRIC_MEAN: frozenset[str]#
‘frozenset(…)’
- nemo_rl.algorithms.single_controller_utils.utils.aggregate_step_metrics(
- train_result: dict[str, Any],
Reduce per-microbatch metric lists into step-level scalars.
- Parameters:
train_result – Output of TQPolicy.finish_train_step.
- Returns:
Flat dict of step-level scalars ready for logging.
- class nemo_rl.algorithms.single_controller_utils.utils.RewardPartial#
One advantage-stage call’s reward sums, already reduced.
Rewards are per-row rather than per-token, so holding the tensors was never the bulk of the controller’s memory. They are reduced here anyway because the advantage stage is moving into its own actor, and the whole point of that boundary is that nothing cohort-sized crosses it.
- weighted_total: float#
None
- weight: float#
None
- classmethod from_rows(
- rewards: torch.Tensor,
- sample_mask: torch.Tensor | None = None,
Reduce one call’s rewards, weighting by row validity when given.
An absent mask weights every row equally, so the merged result is the plain mean and callers that never had a mask keep their old numbers.
- class nemo_rl.algorithms.single_controller_utils.utils.AdvantagePartial#
One advantage-stage call’s token-masked advantage moments.
Replaces keeping
torch.masked_select(advantages, mask)itself, which is one float per trained token across the entire cohort, appended once per streaming chunk and then concatenated at step close – so the step’s peak was twice the accumulated size. Only the mean, min and max were ever read off it, and all three merge from these four numbers.- count: int#
None
- total: float#
None
- minimum: float#
None
- maximum: float#
None
- classmethod from_values(
- values: torch.Tensor,
Reduce one call’s masked advantages; an empty selection counts zero.
- nemo_rl.algorithms.single_controller_utils.utils.reduce_advantage_pump_metrics(
- reward_partials: list[nemo_rl.algorithms.single_controller_utils.utils.RewardPartial],
- advantage_partials: list[nemo_rl.algorithms.single_controller_utils.utils.AdvantagePartial],
- sequence_lengths: list[int],
- *,
- seq_logprob_error_metrics: list[dict[str, float]] | None = None,
- num_mask_sample_filtered: list[int] | None = None,
- num_invalid_tool_calls: list[int] | None = None,
- num_malformed_thinking: list[int] | None = None,
- num_assistant_messages: list[int] | None = None,
- num_routed_experts_backfilled: list[int] | None = None,
Reduce per-step accumulators from _advantage_stage into step scalars.
- Parameters:
reward_partials – One record per advantage_stage call. Already weighted by row validity when the caller had a sample mask (token-capture placeholders and mask_sample/overlong/seq-logprob-error rows carry 0), so
rewardaverages over trained rows only, matching what the advantage estimator’s baseline already excludes.advantage_partials – One record per advantage_stage call, over that call’s token-masked advantages.
sequence_lengths – All input_lengths trained on this step.
seq_logprob_error_metrics – Sequence-error metrics and their aggregation counts, one record per streaming chunk.
num_mask_sample_filtered – Environment-flagged sample counts, one per streaming chunk.
num_invalid_tool_calls – Per-sample invalid tool-call counts.
num_malformed_thinking – Per-sample malformed-thinking counts.
num_assistant_messages – Per-sample assistant message counts (rate denominator).
- Returns:
Step-level reward, advantage, token-count, optional sequence log-probability error metrics, the num_mask_sample_filtered count, and per-sample violation counts.
- nemo_rl.algorithms.single_controller_utils.utils._reduce_seq_logprob_error_metrics(
- records: list[dict[str, float]],
Reduce sequence-error metrics across streaming chunks.
- nemo_rl.algorithms.single_controller_utils.utils.apply_message_level_advantage_penalties(
- advantages: torch.Tensor,
- *,
- invalid_tool_call_mask: torch.Tensor,
- malformed_thinking_mask: torch.Tensor,
- invalid_tool_call_advantage: float | None,
- malformed_thinking_advantage: float | None,
Overwrite flagged token advantages while leaving valid tokens unchanged.
Invalid-tool-call penalties take precedence when both masks select the same token, matching the legacy GRPO message-level implementation.
- Parameters:
advantages – Per-token advantages of shape (batch, seq).
invalid_tool_call_mask – Bool mask of the same shape as
advantages; True at tokens produced by an invalid tool call.malformed_thinking_mask – Bool mask of the same shape as
advantages; True at tokens produced by malformed thinking.invalid_tool_call_advantage – Value to overwrite flagged tokens with, or
Noneto leave the invalid-tool-call branch disabled.malformed_thinking_advantage – Value to overwrite flagged tokens with, or
Noneto leave the malformed-thinking branch disabled.
- Returns:
New tensor with penalties applied. The original
advantagesobject is returned unchanged when both advantages areNone.- Raises:
ValueError – If either mask shape does not match
advantages.
- nemo_rl.algorithms.single_controller_utils.utils.tensor_field(
- data: tensordict.TensorDict,
- field_name: str,
Read a tensor column from a TensorDict, depadding if nested.
- Parameters:
data – TensorDict returned by the data plane.
field_name – Column name to fetch.
- Returns:
Dense tensor (nested columns are padded with zeros).
- nemo_rl.algorithms.single_controller_utils.utils.squeeze_trailing_unit_dim(value: torch.Tensor) torch.Tensor#
Drop a trailing dim of size 1 if present.
- Parameters:
value – Input tensor.
- Returns:
Tensor without the trailing unit dim.
- nemo_rl.algorithms.single_controller_utils.utils.fields_for_put(
- meta: nemo_rl.data_plane.KVBatchMeta,
- fields: dict[str, torch.Tensor],
Pack tensors for DataPlane put, re-nesting jagged rows when needed.
- Parameters:
meta – Batch meta whose sequence_lengths drive the nesting.
fields – Field name to dense tensor.
- Returns:
TensorDict shaped for dp_client.put_samples.