nemo_gym.token_id_capture.terminal

View as Markdown

Attribute a finished rollout’s verified terminal model call.

Every agent’s /run result carries response: the object the verifier scored. Capture, by contrast, records every call the model server served, including auxiliary calls, abandoned retries, and sub-agent branches. Terminal attribution joins the verified response to exactly one captured call, so the builder can deliver the chain that earned the reward instead of guessing among chains or masking a healthy rollout.

The join is independent witnesses with corroboration, not a trust hierarchy. A witness here is one independent reading of the response that can name the captured call. The fact “which call was kept” is emitted in different places by different agent classes, so up to five witnesses testify:

explicit — the caller names the kept call directly (a gate seal or an agent-declared terminal id). Soft: a miss is an abstention and other witnesses may still attribute. declared — the harness reports the response id it retained. Selection stops if that id does not match exactly one captured row. response_id — the /run response’s id equals the served id recorded on exactly one entry. Possession of the id proves which response the client actually received. item_id — the /run response’s last model-authored item carries the id the model server assigned to it (a message’s id, or a tool call’s call_id), and exactly one entry’s recorded output holds that id. A harness that rebuilds its transcript from its own session files keeps these per-item ids even when it never saw the response id, and the last model-authored item names the terminal call even when tool outputs follow it. content — the fingerprint of the response’s model-authored items matches one entry. Three readings are pooled: the entry’s cumulative continuation_fingerprint when a lineage-aware writer recorded one (a full-transcript response), the fingerprint of the entry’s own output (a final-turn-only response), and the transcript’s trailing model-authored block (a merged multi-turn transcript).

External staging rows omit token arrays because those tokens remain in framework storage. They carry a staging_key and precomputed content fingerprints instead. This module identifies external staging rows by the presence of staging_key.

Each witness abstains rather than guesses (ambiguity inside a witness is an abstention, not a vote). The verdict then follows the stack’s rule that claims are verified, never ranked: witnesses that agree — or that name calls with identical full token sequences — attribute; witnesses that contradict each other attribute nothing and persist the disagreement, because a contradiction is evidence of a real defect (a stale seal mapping, backend id reuse, a transcript-synthesis bug) that outranking would silently bury. If no terminal response ID is declared, a rollout with no witness or with disagreeing witnesses falls back to the builder’s strict single-chain policy. If a declared terminal response ID cannot be attributed, the consumer masks the rollout instead.

The content witness deliberately uses assistant_fingerprint alone, without a request-context digest. Attribution selects a chain whose tokens the builder verifies independently; it never reuses tokens across the matched boundary. Requiring context verification would spuriously refuse synthesized transcripts that reformat tool output, without adding safety.

Module Contents

Classes

NameDescription
TerminalAttributionThe joined terminal call, or the reasons no witness could name one.

Functions

NameDescription
_collapse_identicalReduce candidates that carry the same full token sequence to one.
_content_witnessThe content witness: fingerprint the response’s model-authored items.
_entry_items_by_served_idMap every served id in an entry’s recorded output to that item’s content key.
_is_custody_rowA token-free custody row stages its tokens externally under a key.
_item_content_keyReturn the content of one model-authored item in a dialect-independent form.
_item_id_witnessThe item-id witness: the served id of the response’s last model-authored item.
_sequence_identityIdentify an entry by its full token sequence.
_served_item_idReturn the id the model server assigned to an output item, or "" when it carries none.
_tool_name_keyReturn a tool name without the namespace prefix the streaming sanitizer flattens in.
resolve_terminalJoin the verified /run response to one captured model call.

API

class nemo_gym.token_id_capture.terminal.TerminalAttribution(
model_call_id: str | None,
method: str = '',
reason: str = ''
)
Dataclass

The joined terminal call, or the reasons no witness could name one.

attributed
bool
method
str = ''
model_call_id
str | None
reason
str = ''
nemo_gym.token_id_capture.terminal._collapse_identical(

Reduce candidates that carry the same full token sequence to one.

Identical retries produce entries whose sequences match; any of them yields the same delivered chain, so the smallest call id wins deterministically. Candidates with different sequences are genuinely ambiguous and collapse to None.

nemo_gym.token_id_capture.terminal._content_witness(
response: dict,
reasons: list[str]

The content witness: fingerprint the response’s model-authored items.

Three readings of one response are possible and must compete, not race. A full transcript matches an entry’s cumulative fingerprint (the model-authored spine of request context + that call’s output); a final-turn-only response matches the fingerprint of one entry’s own output; a merged transcript’s trailing model-authored block matches the terminal call’s own output. The readings can name different calls — a first call’s cumulative fingerprint IS its own-output fingerprint, because non-model turns never contribute — so candidates from all keys pool before the ambiguity decision.

nemo_gym.token_id_capture.terminal._entry_items_by_served_id(
) -> dict[str, tuple]

Map every served id in an entry’s recorded output to that item’s content key.

The map covers Responses items (a message’s id, a tool call’s call_id) and the tool calls nested in a chat assistant message’s tool_calls list.

nemo_gym.token_id_capture.terminal._is_custody_row(
) -> bool

A token-free custody row stages its tokens externally under a key.

nemo_gym.token_id_capture.terminal._item_content_key(
item: dict
) -> tuple

Return the content of one model-authored item in a dialect-independent form.

A tool call is its name (without any namespace prefix) and canonical arguments; a message is its typed content parts (_content_of reads a plain string and a list of typed parts alike).

nemo_gym.token_id_capture.terminal._item_id_witness(
response: dict,
reasons: list[str]

The item-id witness: the served id of the response’s last model-authored item.

The model server assigns an id to every output item it returns. A harness that rebuilds its transcript from its own session files keeps those ids even when it never records the response id, so the transcript’s last model-authored item names the call that produced it. Tool outputs or user turns that follow it are not model output, so they do not change which item is last.

An id alone does not name a call. The entry’s item with that id must also carry the same content (the tool’s name and canonical arguments, or the message’s content parts), so that a reused or colliding id cannot name a call whose output the transcript does not end with. An item whose id no entry recorded, such as a message the harness composed itself, abstains rather than naming an earlier call.

nemo_gym.token_id_capture.terminal._sequence_identity(
) -> tuple

Identify an entry by its full token sequence.

A custody row identifies by the worker’s whole-sequence cumulative_hash (with the chained chain_hash as a secondary key) plus the cumulative length. A lineage-aware TokenEntry writer records a cumulative digest and length; records without one compare their token arrays directly. All identify the delivered sequence, which is what training consumes.

nemo_gym.token_id_capture.terminal._served_item_id(
item: dict
) -> str

Return the id the model server assigned to an output item, or "" when it carries none.

A tool call is identified by its call_id, the id the serving engine assigned to the call; both the chat and the Responses formats keep that id. Any other item is identified by its id.

nemo_gym.token_id_capture.terminal._tool_name_key(
name: object
) -> str

Return a tool name without the namespace prefix the streaming sanitizer flattens in.

A namespaced Responses tool reaches the chat backend as <namespace>__<name> and is recorded that way in the entry, while the client, and a transcript rebuilt from the client’s files, keeps the bare name with the namespace in a separate field. Two bare names that differ only before their own last __ compare equal here; the witness also requires the served id and the canonical arguments to match, so such a pair cannot be confused on its own.

nemo_gym.token_id_capture.terminal.resolve_terminal(
response: dict | None,
explicit_call_id: str | None = None,
declared_response_id: str | None = None

Join the verified /run response to one captured model call.

entries is the frozen snapshot (TokenEntry records or token-free custody rows). response is the result’s scored response object (or None when the record carries none). This function never raises: malformed content is an abstention, not an error.