nemo_rl.algorithms.single_controller_utils.config#
Module Contents#
Classes#
Fault-tolerance knobs read only by |
|
Fault-tolerance knobs read only by |
|
Fault tolerance for a rollout that fails. |
|
Liveness tracking for the vLLM generation fleet. |
|
NeMo-RL-owned HTTP router placed in front of the vLLM fleet for NeMo-Gym. |
|
Last-resort detection for stalls that no other layer catches. |
|
Ledger-authoritative token capture (token-in/token-out via NeMo-Gym). |
|
Recovery granularity selected for a prompt-group reservation. |
|
Retry and restore policy for unfinished token-capture prompt groups. |
|
Frequent rollout-state snapshots anchored to durable trainer state. |
|
Internal DataPlane field mapping for advantage calculation. |
Functions#
Whether this SingleController run trains a PPO critic alongside the policy. |
|
The active algorithm block: |
|
Validate that backpressure cannot deadlock the selected sampler. |
|
Validate the full-vocabulary MOPD block against the rest of the run. |
|
Check rollout_failure settings that cannot do what they were set for. |
|
Reject algorithm blocks the SingleController path cannot honour. |
|
Validate cross-section SingleController constraints before setup. |
Data#
API#
- class nemo_rl.algorithms.single_controller_utils.config.NativeRolloutFTConfig#
Bases:
pydantic.BaseModelFault-tolerance knobs read only by
AsyncRolloutImpl(the native GRPO path).Setting these on a NeMo-Gym run does nothing;
validate_single_controller_configrejects that rather than letting it pass silently.- generation_timeout_s: Optional[pydantic.PositiveFloat]#
None
- env_timeout_s: Optional[pydantic.PositiveFloat]#
None
- class nemo_rl.algorithms.single_controller_utils.config.NemoGymRolloutFTConfig#
Bases:
pydantic.BaseModelFault-tolerance knobs read only by
AsyncNemoGymRolloutImpl.Setting these on a native run does nothing;
validate_single_controller_configrejects that rather than letting it pass silently.- rollout_timeout_s: Optional[pydantic.PositiveFloat]#
None
- max_row_attempts: pydantic.PositiveInt#
3
- class nemo_rl.algorithms.single_controller_utils.config.RolloutFailureConfig#
Bases:
pydantic.BaseModelFault tolerance for a rollout that fails.
The budgets at the top level are consumed by
generate_and_push, which sits above the native/NeMo-Gym split, so they govern both paths. Everything path-specific lives in thenativeandnemo_gymsub-blocks, so the structure itself says which knob applies where – these used to dangle onasync_rlamong unrelated pump and buffer settings, where a native-path operator could set the most generic-sounding one (rollout_timeout_s) and silently get no deadline at all.Infrastructure failures re-dispatch the prompt onto a different generation shard; data failures are deterministic, so their budget is small and exhausting it is reported rather than absorbed. Nothing here ever discards a prompt silently.
Each class also has a budget for how many prompts may be given up on entirely, and the two differ because the question they answer differs. Data exhaustion is a property of the dataset, so
max_skipped_promptscounts them for the run’s lifetime. Infra exhaustion is a property of the fleet at a moment in time, somax_consecutive_dropped_promptsresets on every success: an outage that ends is absorbed, one that does not stops the run.Once a prompt has been given up on,
on_dropped_promptdecides what happens to the training step it was stamped for: train the step on fewer groups, or substitute a fresh prompt so the step keeps its configured batch size.- max_infra_attempts_per_prompt: pydantic.PositiveInt#
5
- max_data_attempts_per_prompt: pydantic.PositiveInt#
2
- backoff_base_s: pydantic.PositiveFloat#
1.0
- max_backoff_s: pydantic.PositiveFloat#
30.0
- max_skipped_prompts: pydantic.NonNegativeInt#
0
- max_consecutive_dropped_prompts: pydantic.NonNegativeInt#
0
- min_step_batch_fraction: Annotated[float, Field(gt=0.0, le=1.0)]#
0.9
- on_dropped_prompt: Literal[shrink, replace]#
‘shrink’
- max_replacement_attempts: pydantic.NonNegativeInt#
1
- replacement_reserve_prompts: pydantic.NonNegativeInt#
1
- native: nemo_rl.algorithms.single_controller_utils.config.NativeRolloutFTConfig#
‘Field(…)’
- nemo_gym: nemo_rl.algorithms.single_controller_utils.config.NemoGymRolloutFTConfig#
‘Field(…)’
- _check_consistent() nemo_rl.algorithms.single_controller_utils.config.RolloutFailureConfig#
- _reject_renamed_keys() nemo_rl.algorithms.single_controller_utils.config.RolloutFailureConfig#
Fail loudly on the previous key names rather than ignoring them.
extra="allow"means an old key parses fine and then does nothing. Foron_data_exhausted: skipthat is a behaviour change – prompts that used to be skipped now fail the run – arriving with no diagnostic at all.
- class nemo_rl.algorithms.single_controller_utils.config.FleetHealthConfig#
Bases:
pydantic.BaseModelLiveness tracking for the vLLM generation fleet.
Only the knobs P1 actually consumes are declared. Recovery modes beyond
fail_fastneed the communicator rebuild that lands later, so the Literal rejects them rather than accepting a value that would silently do nothing.- enabled: bool#
False
- probe_interval_s: pydantic.PositiveFloat#
5.0
- probe_timeout_s: pydantic.PositiveFloat#
2.0
- unhealthy_threshold: pydantic.PositiveInt#
3
- healthy_threshold: pydantic.PositiveInt#
2
- selection: Literal[least_outstanding]#
‘least_outstanding’
- max_restart_attempts_per_shard: pydantic.PositiveInt#
5
- min_healthy_shards: pydantic.PositiveInt#
1
- refit_timeout_s: Optional[pydantic.PositiveFloat]#
300.0
- restart_dead_shards: bool#
False
- restart_timeout_s: pydantic.PositiveFloat#
1800.0
- restart_backoff_s: pydantic.PositiveFloat#
60.0
- _check_consistent() nemo_rl.algorithms.single_controller_utils.config.FleetHealthConfig#
- nemo_rl.algorithms.single_controller_utils.config._GYM_RETRY_STATUSES: frozenset[int]#
‘frozenset(…)’
- class nemo_rl.algorithms.single_controller_utils.config.GenerationRouterConfig#
Bases:
pydantic.BaseModelNeMo-RL-owned HTTP router placed in front of the vLLM fleet for NeMo-Gym.
Gym selects a policy endpoint by static round-robin over a list fixed at process start and never fails over. Handing it a single NeMo-RL-owned URL moves that decision to where fleet health already lives, without changing Gym.
- enabled: bool#
False
- port_range_low: pydantic.PositiveInt#
None
- port_range_high: pydantic.PositiveInt#
None
- backend_timeout_s: pydantic.PositiveFloat#
600.0
- connect_timeout_s: pydantic.PositiveFloat#
5.0
- no_healthy_backend_status: pydantic.PositiveInt#
409
- _check_port_range() nemo_rl.algorithms.single_controller_utils.config.GenerationRouterConfig#
- _check_connect_timeout_fits() nemo_rl.algorithms.single_controller_utils.config.GenerationRouterConfig#
- _check_status_is_not_retried_by_gym() nemo_rl.algorithms.single_controller_utils.config.GenerationRouterConfig#
- class nemo_rl.algorithms.single_controller_utils.config.WatchdogConfig#
Bases:
pydantic.BaseModelLast-resort detection for stalls that no other layer catches.
- interval_s: pydantic.PositiveFloat#
30.0
- stall_timeout_s: pydantic.PositiveFloat#
600.0
- stall_action: Literal[warn, abort]#
‘warn’
- gym_subprocess_check: bool#
True
- _check_consistent() nemo_rl.algorithms.single_controller_utils.config.WatchdogConfig#
- class nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig#
Bases:
pydantic.BaseModel- log_full_train_data: bool#
False
- sampler: nemo_rl.algorithms.async_utils.staleness_sampler.SamplerConfig#
‘Field(…)’
- rollout_failure: nemo_rl.algorithms.single_controller_utils.config.RolloutFailureConfig#
‘Field(…)’
- stall_watchdog: nemo_rl.algorithms.single_controller_utils.config.WatchdogConfig#
‘Field(…)’
- generation_fleet_health: nemo_rl.algorithms.single_controller_utils.config.FleetHealthConfig#
‘Field(…)’
- generation_router: nemo_rl.algorithms.single_controller_utils.config.GenerationRouterConfig#
‘Field(…)’
- recompute_kv_cache_after_weight_updates: bool#
False
- min_groups_for_streaming_train: int#
32
- max_inflight_prompts: int#
32
- max_buffered_rollouts: int#
64
- diagnostics: bool#
False
- num_advantage_workers: pydantic.NonNegativeInt#
0
- _reject_renamed_blocks() nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig#
Fail loudly on the previous block names rather than ignoring them.
extra="allow"means an old key parses fine and then does nothing at all – so a config carryingwatchdog:would silently lose its stall detection and run with the defaults, which is precisely the class of silent misconfiguration this work exists to remove.watchdogin particular shipped, so this is a migration path rather than a courtesy.
- _check_router_deadline_fits_inside_the_rollout() nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig#
The router’s per-request deadline must not outlast the whole rollout’s.
backend_timeout_sbounds ONE HTTP call;rollout_timeout_sbounds the whole prompt-group stream, which is many of them. Set the inner one larger and it can never fire: the rollout deadline always expires first, so the timeout the router exists to add is dead config – the silent no-op shape this series exists to remove. It is also the wrong failure to surface, because the rollout layer reports the group while the router could have named the backend.
- _check_stall_watchdog_outlasts_rollouts() nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig#
- _reject_relocated_keys() nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig#
Fail loudly on keys that moved, instead of ignoring them.
These models are
extra="allow", so a config written against the previous layout keeps parsing and its fault-tolerance settings simply stop taking effect. A silently ignoredrollout_timeout_s: 900is precisely the failure mode the restructure was meant to remove, so the move must not create one on its way out.
- class nemo_rl.algorithms.single_controller_utils.config.TokenCaptureConfig#
Bases:
pydantic.BaseModelLedger-authoritative token capture (token-in/token-out via NeMo-Gym).
Dormant by default: with
enabled=Falseevery legacy codepath behaves exactly as before — no staging partition is registered, no ledger is installed, and rollouts ride the token-echo path.- enabled: bool#
False
- staging_partition: str#
‘rollout_staging’
- min_valid_fraction_per_group: Optional[float]#
None
- control_auth_token: Optional[str]#
None
- control_timeout_s: float#
60.0
- capture_dir: Optional[str]#
None
- generation_backend: Optional[Literal[vllm, megatron]]#
None
- defer_routed_experts_to_policy: bool#
False
- num_reassembler_workers: pydantic.PositiveInt#
2
- class nemo_rl.algorithms.single_controller_utils.config.TaskSourceRecoveryGranularity#
Recovery granularity selected for a prompt-group reservation.
task_sourceis copied from the raw Gym row when present.granularityis selected from an explicit agent override, a task-source override, or the global default.- task_source: Optional[str]#
None
- granularity: nemo_rl.experience.rollout_recovery.RecoveryGranularity#
None
- class nemo_rl.algorithms.single_controller_utils.config.RolloutRecoveryConfig#
Bases:
pydantic.BaseModelRetry and restore policy for unfinished token-capture prompt groups.
sibling(the default) preserves completed generations and retries only the missing ones. Prefer it when reusing work and avoiding repeated long-tail generations matters more than keeping a group on one policy version.prompt_groupdiscards and regenerates every sibling when any generation is unfinished. It costs a full group per recovery, but keeps the regenerated group on the policy weights live at redispatch instead of mixing those results with older sealed siblings.The resolved value is persisted on each ledger group, so restoring a saved group does not reinterpret it using a newer configuration. The same granularity governs failures handled in-process and after a process restart.
- default_granularity: nemo_rl.experience.rollout_recovery.RecoveryGranularity#
None
- task_source_granularity_overrides: dict[str, nemo_rl.experience.rollout_recovery.RecoveryGranularity]#
‘Field(…)’
- agent_granularity_overrides: dict[str, nemo_rl.experience.rollout_recovery.RecoveryGranularity]#
‘Field(…)’
- _reject_removed_override_keys() nemo_rl.algorithms.single_controller_utils.config.RolloutRecoveryConfig#
Reject the removed task-name map instead of silently ignoring it.
- resolve_for_prompt(
- prompt: collections.abc.Mapping[str, Any],
Resolve using matching agent, matching task source, then default.
- class nemo_rl.algorithms.single_controller_utils.config.RolloutCheckpointConfig#
Bases:
pydantic.BaseModelFrequent rollout-state snapshots anchored to durable trainer state.
snapshot_attempt_interval_s=Nonedisables saving and restoring periodic snapshots. A snapshot taken before the first trainer checkpoint is anchored to the initial model and a rollout-semantic configuration fingerprint. Later snapshots require the durable trainer checkpoint for the controller’s current completed step; attempts are skipped until that exact anchor exists.restore_mode="latest"selects the newest compatible periodic snapshot.trainer_checkpointignores newer periodic snapshots and restores the rollout state bundled with the durable trainer checkpoint. Restore selection never deletes checkpoint state. If no trainer checkpoint exists,trainer_checkpointrejects an occupied bootstrap namespace; uselatestor a new checkpoint directory instead.Bootstrap compatibility is fail-closed: every configuration value affects the fingerprint unless it is on the built-in operational denylist.
extra_fingerprint_excluded_pathslets integrations exclude additional runtime-only dotpaths.*matches one mapping or list level and**matches any number of levels.SingleController has no validation loop, so checkpoint selection must use
checkpointing.metric_name=Noneor atrain:<name>metric. Inheritedval:<name>settings are rejected during setup. Unknown keys are forbidden because a misspelled interval, retention, or restore option can silently disable the durability behavior the operator intended.telemetry_interval_s=Nonedisables the independent wall-clock sampler for rollout/checkpoint benchmark metrics. It does not enable checkpointing and may be configured withoutsnapshot_attempt_interval_s.max_consecutive_failurescontrols how many consecutive retryable periodic-checkpoint failures are tolerated before the controller aborts the run. A successful or skipped attempt resets the counter; checkpoint invariant failures still fail immediately.- snapshot_attempt_interval_s: Annotated[Optional[float], Field(gt=0)]#
None
- telemetry_interval_s: Annotated[Optional[float], Field(gt=0)]#
None
- max_consecutive_failures: Annotated[int, Field(ge=1)]#
3
- keep_latest_k: Annotated[int, Field(ge=1)]#
2
- restore_mode: Literal[latest, trainer_checkpoint]#
‘latest’
- extra_fingerprint_excluded_paths: list[str]#
‘Field(…)’
- validate_extra_fingerprint_excluded_paths() nemo_rl.algorithms.single_controller_utils.config.RolloutCheckpointConfig#
Reject ambiguous paths that could silently fail to exclude a value.
- class nemo_rl.algorithms.single_controller_utils.config.MasterConfig#
Bases:
pydantic.BaseModel- grpo: Optional[nemo_rl.algorithms.grpo.GRPOConfig]#
None
- ppo: Optional[nemo_rl.algorithms.ppo.PPOConfig]#
None
- policy: nemo_rl.models.policy.PolicyConfig#
None
- value: Optional[nemo_rl.models.value.ValueConfig]#
None
- loss_fn: nemo_rl.algorithms.loss.ClippedPGLossConfig#
None
- value_loss_fn: Optional[nemo_rl.algorithms.loss.loss_functions.MseValueLossConfig]#
None
- env: dict[str, Any]#
None
- data: nemo_rl.data.DataConfig#
None
- logger: nemo_rl.utils.logger.LoggerConfig#
None
- cluster: nemo_rl.distributed.virtual_cluster.ClusterConfig#
None
- checkpointing: nemo_rl.utils.checkpoint.CheckpointingConfig#
None
- reward_penalties: nemo_rl.algorithms.grpo.RewardPenaltyConfig#
‘Field(…)’
- data_plane: nemo_rl.data_plane.interfaces.DataPlaneConfig#
None
- rollout_recovery: nemo_rl.algorithms.single_controller_utils.config.RolloutRecoveryConfig#
‘Field(…)’
- rollout_checkpointing: nemo_rl.algorithms.single_controller_utils.config.RolloutCheckpointConfig#
‘Field(…)’
- on_policy_distillation: Optional[nemo_rl.algorithms.opd.OnPolicyDistillationConfig]#
None
- telemetry: Optional[nemo_rl.telemetry.config.TelemetryConfig]#
None
- token_capture: nemo_rl.algorithms.single_controller_utils.config.TokenCaptureConfig#
‘Field(…)’
- validate_algorithm_block() nemo_rl.algorithms.single_controller_utils.config.MasterConfig#
- nemo_rl.algorithms.single_controller_utils.config.is_ppo_run(
- master_config: nemo_rl.algorithms.single_controller_utils.config.MasterConfig,
Whether this SingleController run trains a PPO critic alongside the policy.
Single source of truth for the flag: setup reads it to decide whether to build the value model, and the controller reads it to decide whether the train pump runs the critic stages.
model_constructskips defaults, so the attribute can genuinely be missing on a hand-built config.
- nemo_rl.algorithms.single_controller_utils.config.algo_config(
- master_config: nemo_rl.algorithms.single_controller_utils.config.MasterConfig,
The active algorithm block:
ppoon a PPO run, elsegrpo.Exactly one of the two is set; MasterConfig.validate_algorithm_block checks that at construction.
- nemo_rl.algorithms.single_controller_utils.config.validate_sampler_buffer_capacity(
- async_config: nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig,
- *,
- required_capacity: Optional[int],
- sampler_name: str,
Validate that backpressure cannot deadlock the selected sampler.
- nemo_rl.algorithms.single_controller_utils.config._validate_opd_full_config(
- master_config: nemo_rl.algorithms.single_controller_utils.config.MasterConfig,
- opd_config: nemo_rl.algorithms.opd.OnPolicyDistillationConfig,
Validate the full-vocabulary MOPD block against the rest of the run.
- Parameters:
master_config – Full SingleController config, already known to have OPD on.
opd_config – The resolved
on_policy_distillationblock.
- Raises:
ValueError – If
opd_fullis enabled with an unsupported backend, an incompatible logprob path, a fused packing path that never reaches the opd_full branch, or a sampling temperature the hidden-state payload cannot honor.
- nemo_rl.algorithms.single_controller_utils.config._validate_failure_settings(
- async_config: nemo_rl.algorithms.single_controller_utils.config.AsyncRLConfig,
- num_prompts_per_step: int,
Check rollout_failure settings that cannot do what they were set for.
RolloutFailureConfig._check_consistentrejects the two combinations that makeon_dropped_prompt="replace"unable to ever produce a replacement. These are the combinations it cannot see, because each depends on a field outside that block.Warnings rather than errors for most of them, deliberately: a strict “never train a short batch” run, one base YAML whose sampler is overridden per experiment, a deliberately deep spare pool are all coherent things to have typed, so rejecting them would forbid configurations somebody wants. What is not acceptable is finding out hours into a run, or not at all.
The exception, and the only hard error here, is a drop budget under a sampler that stamps no target step. That one is not a knob cancelling itself out; it converts a recoverable prompt failure into a stalled run, so there is no configuration it could be the intent of.
- nemo_rl.algorithms.single_controller_utils.config._validate_algo_settings(
- master_config: nemo_rl.algorithms.single_controller_utils.config.MasterConfig,
Reject algorithm blocks the SingleController path cannot honour.
Both directions on the critic: one the PPO path needs and does not have, and one a GRPO run carries and would never build. Plus the reward-shaping and sampling knobs SC reads on neither path.
- nemo_rl.algorithms.single_controller_utils.config.validate_single_controller_config(
- master_config: nemo_rl.algorithms.single_controller_utils.config.MasterConfig,
Validate cross-section SingleController constraints before setup.
- class nemo_rl.algorithms.single_controller_utils.config.AdvantageConfig#
Internal DataPlane field mapping for advantage calculation.
- output_field: str#
‘advantages’
- reward_field: str#
‘total_reward’
- token_mask_field: str#
‘token_mask’
- sample_mask_field: str#
‘sample_mask’
- invalid_tool_call_mask_field: str#
None
- malformed_thinking_mask_field: str#
None
- mask_sample_field: str#
‘mask_sample’
- truncated_field: str#
‘truncated’
- repeated_batch_fields: list[str]#
‘field(…)’
- policy_logprobs_field: str#
‘prev_logprobs’
- generation_logprobs_field: str#
‘generation_logprobs’
- reference_logprobs_field: str#
‘reference_policy_logprobs’
- teacher_logprobs_field: str#
‘teacher_reference_logprobs’
- values_field: str#
‘values’
- returns_field: str#
‘returns’
- prompt_ids_field: str#
‘prompt_ids_for_adv’