nemo_gym.health.checks

View as Markdown

Artifact normalization and single-rollout health checks.

Module Contents

Functions

NameDescription
_agent_stepsNormalize canonical TrajectoryTurn records for structural checks.
_agent_turn_hollow-
_bind_policy_call_viewsBind turn-only and turn-or-invocation references with shared indexes and resolutions.
_call_identity-
_call_locator-
_call_ref_key-
_call_reference_signature-
_canonical_trajectory-
_chain_ended_on_failureWhether this model’s own last call, in time, failed.
_deduplicate_reference_items-
_ended_on_failed_callWhether any model’s own last call failed.
_is_failed-
_is_successful-
_item_has_tool_call-
_item_model_call_referencesReturn usable model-call references from turns or invocations.
_model_call_failed-
_model_call_missing_token_counts-
_model_call_runaway_generation-
_model_call_zero_completion_tokens-
_model_chainsGroup calls by the model they went to.
_nonempty-
_normalized_trajectory_calls-
_replay_identity-
_response_has_content-
_rollout_ended_on_failed_model_callFlag a model whose own last call in the rollout failed.
_rollout_missing_agent_turns-
_rollout_token_count_mismatch-
_subject-
_token_count-
_trajectory_capture_mismatch-
_trajectory_has_any_gap-
_trajectory_has_gap-
_trajectory_reference_contradictions-
_transcript_tokens-
_usage_tokens-
normalize_ignored_checksNormalize and validate check IDs supplied by library, CLI, or Hydra config.

Data

CHECK_REGISTRY

_AGENT_TOOL_CALL_TYPES

_INCOMPLETE_MODEL_CALL_GAPS

_LENGTH_LIMIT_FINISH_REASONS

_REFERENCE_CONTRADICTION_GAPS

_ROLLOUT_CHECKS

_ROLLOUT_SPECS

_TASK_SPECS

API

nemo_gym.health.checks._agent_steps(
trajectory: dict[str, typing.Any]

Normalize canonical TrajectoryTurn records for structural checks.

nemo_gym.health.checks._agent_turn_hollow(
trajectory: dict[str, typing.Any],
subject: dict[str, int | str]
nemo_gym.health.checks._bind_policy_call_views(
trajectory: dict[str, typing.Any],
calls: list[dict[str, typing.Any]]

Bind turn-only and turn-or-invocation references with shared indexes and resolutions.

nemo_gym.health.checks._call_identity(
call: dict[str, typing.Any]
) -> str | None
nemo_gym.health.checks._call_locator(
call: dict[str, typing.Any],
fallback: int
) -> dict[str, int | str]
nemo_gym.health.checks._call_ref_key(
ref: typing.Any
) -> str | None
nemo_gym.health.checks._call_reference_signature(
reference: dict[str, typing.Any]
) -> tuple[str, str, str, str]
nemo_gym.health.checks._canonical_trajectory(
record: dict[str, typing.Any]
) -> tuple[dict[str, typing.Any] | None, str | None]
nemo_gym.health.checks._chain_ended_on_failure(
chain: collections.abc.Sequence[dict[str, typing.Any]]
) -> bool

Whether this model’s own last call, in time, failed.

Per model rather than across every captured call: a judge or auxiliary request landing after a failed policy call would otherwise read as the rollout recovering, and hide the failure.

Without timestamps a multi-call chain cannot be ordered — call_index is the position in the merged trajectory list, so a producer-supplied retry can precede the attempt it retried. Say nothing rather than call a recovered rollout unhealthy.

nemo_gym.health.checks._deduplicate_reference_items(
reference_items: collections.abc.Sequence[tuple[str, dict[str, typing.Any]]]
) -> tuple[tuple[str, dict[str, typing.Any]], ...]
nemo_gym.health.checks._ended_on_failed_call(
calls: collections.abc.Sequence[dict[str, typing.Any]]
) -> bool

Whether any model’s own last call failed.

nemo_gym.health.checks._is_failed(
call: dict[str, typing.Any]
) -> bool
nemo_gym.health.checks._is_successful(
call: dict[str, typing.Any]
) -> bool
nemo_gym.health.checks._item_has_tool_call(
item: typing.Any
) -> bool
nemo_gym.health.checks._item_model_call_references(
items: typing.Any
) -> tuple[tuple[str, dict[str, typing.Any]], ...]

Return usable model-call references from turns or invocations.

nemo_gym.health.checks._model_call_failed(
subject: dict[str, int | str]
nemo_gym.health.checks._model_call_missing_token_counts(
subject: dict[str, int | str]
nemo_gym.health.checks._model_call_runaway_generation(
subject: dict[str, int | str]
nemo_gym.health.checks._model_call_zero_completion_tokens(
subject: dict[str, int | str]
nemo_gym.health.checks._model_chains(
calls: collections.abc.Sequence[dict[str, typing.Any]]
) -> dict[typing.Any, list[dict[str, typing.Any]]]

Group calls by the model they went to.

nemo_gym.health.checks._nonempty(
value: typing.Any
) -> bool
nemo_gym.health.checks._normalized_trajectory_calls(
trajectory: dict[str, typing.Any]
) -> list[dict[str, typing.Any]]
nemo_gym.health.checks._replay_identity(
call: dict[str, typing.Any]
) -> str | None
nemo_gym.health.checks._response_has_content(
response: typing.Any
) -> bool
nemo_gym.health.checks._rollout_ended_on_failed_model_call(
trajectory: dict[str, typing.Any],
subject: dict[str, int | str]

Flag a model whose own last call in the rollout failed.

The bound-call checks cannot see this: binding resolves a reference by (model_ref, response_id) or model_call_id, and a call that failed came back with none of them, so it is absent from matched_calls no matter what the producer claims. Reading the captured calls directly also covers the agents that publish no trajectory at all.

Judged per model and in time order — see _chain_ended_on_failure. Only the chain’s last call counts, so a failure the client retried successfully stays healthy; the signal is that the rollout ENDED on a failure, which is what makes its reward indistinguishable from a genuine zero.

nemo_gym.health.checks._rollout_missing_agent_turns(
trajectory: dict[str, typing.Any],
subject: dict[str, int | str]
nemo_gym.health.checks._rollout_token_count_mismatch(
record: dict[str, typing.Any],
subject: dict[str, int | str]
nemo_gym.health.checks._subject(
task_index: int | str,
rollout_index: int | str | None = None
) -> dict[str, int | str]
nemo_gym.health.checks._token_count(
call: dict[str, typing.Any],
key: str
) -> int
nemo_gym.health.checks._trajectory_capture_mismatch(
trajectory: dict[str, typing.Any],
subject: dict[str, int | str]
nemo_gym.health.checks._trajectory_has_any_gap(
trajectory: dict[str, typing.Any],
codes: frozenset[str]
) -> bool
nemo_gym.health.checks._trajectory_has_gap(
trajectory: dict[str, typing.Any],
code: str
) -> bool
nemo_gym.health.checks._trajectory_reference_contradictions(
trajectory: dict[str, typing.Any]
) -> list[dict[str, typing.Any]]
nemo_gym.health.checks._transcript_tokens(
record: dict[str, typing.Any]
) -> tuple[int, int, bool]
nemo_gym.health.checks._usage_tokens(
usage: typing.Any
) -> tuple[int | None, int | None]
nemo_gym.rollout_health.normalize_ignored_checks(
checks: collections.abc.Sequence[str] | str | None
) -> tuple[str, ...]

Normalize and validate check IDs supplied by library, CLI, or Hydra config.

nemo_gym.health.checks.CHECK_REGISTRY: tuple[CheckSpec, ...] = (CheckSpec(id='check_execution_error', evaluation_scope=(CheckScope.ROLLOUT), su...
nemo_gym.health.checks._AGENT_TOOL_CALL_TYPES = frozenset({'function_call', 'tool_call', 'tool_use', 'mcp_call', 'mcp_list_tools...
nemo_gym.health.checks._INCOMPLETE_MODEL_CALL_GAPS = frozenset({'model_calls_unavailable', 'model_call_capture_incomplete', 'model_ca...
nemo_gym.health.checks._LENGTH_LIMIT_FINISH_REASONS = frozenset({'length', 'max_output_tokens', 'max_tokens'})
nemo_gym.health.checks._REFERENCE_CONTRADICTION_GAPS = {'model_call_reference_unmatched': 'missing_captured_call', 'model_call_referenc...
nemo_gym.health.checks._ROLLOUT_CHECKS: dict[str, Callable[[dict[str, Any], dict[str, Any], _CallBindings, dict[str, int | str]], list[Finding]]] = {'check_execution_error': lambda record, trajectory, bindings, subject: [], 'rec...
nemo_gym.health.checks._ROLLOUT_SPECS = tuple(spec for spec in CHECK_REGISTRY if spec.evaluation_scope == CheckScope.ROL...
nemo_gym.health.checks._TASK_SPECS = tuple(spec for spec in CHECK_REGISTRY if spec.evaluation_scope == CheckScope.TAS...