> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemo/gym/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemo/gym/_mcp/server.

# nemo_gym.base_responses_api_model

Model server base classes and per-rollout model-call capture.

Every Gym model server derives from `SimpleResponsesAPIModel`, which wires the three model
dialects (/v1/responses, /v1/chat/completions, /v1/messages) and installs the model-call capture
middleware.

Capture is opt-in, off by default. A pure-ASGI middleware records correlated /v1/responses,
/v1/chat/completions, and /v1/messages exchanges -- including failed calls -- into a
per-rollout CaptureStore, forwarding bytes downstream unchanged so it composes with
streaming (SSE) responses. Best-effort; never alters the response. Correlation is
carried by a /ng-rollout/\<rollout\_id>/v1/... base\_url prefix, which is stripped before
routing.

## Module Contents

### Classes

| Name                                                                                            | Description                                                           |
| ----------------------------------------------------------------------------------------------- | --------------------------------------------------------------------- |
| [`BaseResponsesAPIModel`](#nemo_gym-base_responses_api_model-BaseResponsesAPIModel)             | -                                                                     |
| [`BaseResponsesAPIModelConfig`](#nemo_gym-base_responses_api_model-BaseResponsesAPIModelConfig) | -                                                                     |
| [`CaptureStore`](#nemo_gym-base_responses_api_model-CaptureStore)                               | Append-only, rollout-keyed JSONL sink for model exchanges.            |
| [`ModelCallCaptureConfig`](#nemo_gym-base_responses_api_model-ModelCallCaptureConfig)           | Run-wide model-call capture settings from Gym's global config.        |
| [`ModelCallRecord`](#nemo_gym-base_responses_api_model-ModelCallRecord)                         | Observability record derived from one captured model-server exchange. |
| [`ModelExecutionOutcome`](#nemo_gym-base_responses_api_model-ModelExecutionOutcome)             | -                                                                     |
| [`SimpleResponsesAPIModel`](#nemo_gym-base_responses_api_model-SimpleResponsesAPIModel)         | -                                                                     |
| [`_CaptureMiddleware`](#nemo_gym-base_responses_api_model-_CaptureMiddleware)                   | Pure-ASGI per-rollout capture.                                        |

### Functions

| Name                                                                                                                  | Description                                                                                    |
| --------------------------------------------------------------------------------------------------------------------- | ---------------------------------------------------------------------------------------------- |
| [`_as_arguments`](#nemo_gym-base_responses_api_model-_as_arguments)                                                   | -                                                                                              |
| [`_cache_signal`](#nemo_gym-base_responses_api_model-_cache_signal)                                                   | Cache hit/miss + cached-token count, from usage cache fields (OpenAI / Anthropic).             |
| [`_classify_exception`](#nemo_gym-base_responses_api_model-_classify_exception)                                       | Normalized error\_category for an exception raised while calling the model.                    |
| [`_classify_status`](#nemo_gym-base_responses_api_model-_classify_status)                                             | Normalized error\_category from an HTTP status (None when \< 400).                             |
| [`_consume_terminal_sse_event`](#nemo_gym-base_responses_api_model-_consume_terminal_sse_event)                       | -                                                                                              |
| [`_exception_http_details`](#nemo_gym-base_responses_api_model-_exception_http_details)                               | -                                                                                              |
| [`_fail_uncommitted_external_call`](#nemo_gym-base_responses_api_model-_fail_uncommitted_external_call)               | Record a failure when an admitted worker call returns without commit coordinates.              |
| [`_headers_content_type`](#nemo_gym-base_responses_api_model-_headers_content_type)                                   | -                                                                                              |
| [`_orjson_dispatch_response`](#nemo_gym-base_responses_api_model-_orjson_dispatch_response)                           | Serialize a completed non-streaming model response with orjson.                                |
| [`_parse_sse_events`](#nemo_gym-base_responses_api_model-_parse_sse_events)                                           | Parse an SSE byte stream into its JSON `data:` payloads (best-effort; non-JSON skipped).       |
| [`_plain`](#nemo_gym-base_responses_api_model-_plain)                                                                 | Reduce a request field to plain data so equal schemas serialize equally.                       |
| [`_preserve_capture_prefix_on_redirect`](#nemo_gym-base_responses_api_model-_preserve_capture_prefix_on_redirect)     | Keep rollout correlation on root-relative and same-origin redirects.                           |
| [`_reconstruct_anthropic_sse`](#nemo_gym-base_responses_api_model-_reconstruct_anthropic_sse)                         | Rebuild a complete Anthropic Messages response from its streamed events.                       |
| [`_reconstruct_chat_sse`](#nemo_gym-base_responses_api_model-_reconstruct_chat_sse)                                   | Rebuild a Chat Completions response from streamed chunks.                                      |
| [`_reconstruct_responses_sse`](#nemo_gym-base_responses_api_model-_reconstruct_responses_sse)                         | Rebuild a Responses API response: the terminal envelope carries the full response object.      |
| [`_reconstruct_streamed_response`](#nemo_gym-base_responses_api_model-_reconstruct_streamed_response)                 | Best-effort: reassemble a final response object from a streamed (SSE) body, by dialect.        |
| [`_record`](#nemo_gym-base_responses_api_model-_record)                                                               | Append one exchange (success or failure). Best-effort: never raises.                           |
| [`_request_envelope`](#nemo_gym-base_responses_api_model-_request_envelope)                                           | Return prompt-shaping request fields that are not turns.                                       |
| [`_request_messages`](#nemo_gym-base_responses_api_model-_request_messages)                                           | Return the conversation carried by any supported dialect.                                      |
| [`_store_for_rollout`](#nemo_gym-base_responses_api_model-_store_for_rollout)                                         | -                                                                                              |
| [`_token_count`](#nemo_gym-base_responses_api_model-_token_count)                                                     | -                                                                                              |
| [`_tool_calls_and_reasoning`](#nemo_gym-base_responses_api_model-_tool_calls_and_reasoning)                           | Structured tool calls (name, arguments, call\_id) and reasoning text, across all three shapes. |
| [`_unique_request_header`](#nemo_gym-base_responses_api_model-_unique_request_header)                                 | -                                                                                              |
| [`_usage_detail_token`](#nemo_gym-base_responses_api_model-_usage_detail_token)                                       | Return the first valid token count across equivalent provider detail shapes.                   |
| [`_validate_chat_params`](#nemo_gym-base_responses_api_model-_validate_chat_params)                                   | Validate a /v1/chat/completions body dict, surfacing failures as FastAPI's standard 422.       |
| [`_validate_responses_params`](#nemo_gym-base_responses_api_model-_validate_responses_params)                         | Validate a /v1/responses body dict, surfacing failures as FastAPI's standard 422.              |
| [`_validate_rollout_id`](#nemo_gym-base_responses_api_model-_validate_rollout_id)                                     | -                                                                                              |
| [`aggregate_model_call_metrics`](#nemo_gym-base_responses_api_model-aggregate_model_call_metrics)                     | Aggregate model-call metrics for one rollout id.                                               |
| [`aggregate_model_call_records`](#nemo_gym-base_responses_api_model-aggregate_model_call_records)                     | Aggregate token and latency values from model-call records.                                    |
| [`build_model_call_record`](#nemo_gym-base_responses_api_model-build_model_call_record)                               | Map one captured exchange and its transport metadata into an observability record.             |
| [`clear_model_call_captures_for_rollouts`](#nemo_gym-base_responses_api_model-clear_model_call_captures_for_rollouts) | Remove stale per-rollout capture files for these records before dispatch.                      |
| [`extract_token_stats`](#nemo_gym-base_responses_api_model-extract_token_stats)                                       | Normalize token totals across Responses, Chat Completions, and Anthropic Messages usage.       |
| [`install_model_call_capture`](#nemo_gym-base_responses_api_model-install_model_call_capture)                         | Install model-call capture middleware.                                                         |
| [`make_capture_store`](#nemo_gym-base_responses_api_model-make_capture_store)                                         | Build a CaptureStore when observability is enabled; otherwise None.                            |
| [`merge_model_call_capture_into_record`](#nemo_gym-base_responses_api_model-merge_model_call_capture_into_record)     | Attach captured model-call observability data to a rollout record in place.                    |
| [`model_call_capture_dirs_from_config`](#nemo_gym-base_responses_api_model-model_call_capture_dirs_from_config)       | Return the single run-wide capture directory when capture is enabled.                          |
| [`observability_enabled_from_config`](#nemo_gym-base_responses_api_model-observability_enabled_from_config)           | Return the run-wide `observability_enabled` flag directly, without going through capture dirs. |
| [`read_available_model_call_records`](#nemo_gym-base_responses_api_model-read_available_model_call_records)           | Read valid call records and count damaged records.                                             |
| [`read_model_call_records`](#nemo_gym-base_responses_api_model-read_model_call_records)                               | Read captured exchanges in durable append order.                                               |
| [`start_model_execution`](#nemo_gym-base_responses_api_model-start_model_execution)                                   | Keep adapter-owned execution facts separate from the served response and capture settings.     |

### Data

[`_ANTHROPIC_CONVERTER`](#nemo_gym-base_responses_api_model-_ANTHROPIC_CONVERTER)

[`_CHAT_KEEPALIVE_SECONDS`](#nemo_gym-base_responses_api_model-_CHAT_KEEPALIVE_SECONDS)

[`_CLIENT_SESSION_HEADER`](#nemo_gym-base_responses_api_model-_CLIENT_SESSION_HEADER)

[`_ENVELOPE_ROLE`](#nemo_gym-base_responses_api_model-_ENVELOPE_ROLE)

[`_OBSERVED_PATHS`](#nemo_gym-base_responses_api_model-_OBSERVED_PATHS)

[`_ROLLOUT_PATH_RE`](#nemo_gym-base_responses_api_model-_ROLLOUT_PATH_RE)

[`_SSE_KEEPALIVE`](#nemo_gym-base_responses_api_model-_SSE_KEEPALIVE)

[`_TERMINAL_SSE_LINES`](#nemo_gym-base_responses_api_model-_TERMINAL_SSE_LINES)

[`logger`](#nemo_gym-base_responses_api_model-logger)

### API

```python
class nemo_gym.base_responses_api_model.BaseResponsesAPIModel()
```

**Bases:** [BaseServer](/nemo/gym/nemo-gym/nemo_gym/server_utils#nemo_gym-server_utils-BaseServer)

**`config`** `BaseResponsesAPIModelConfig`

---

```python
class nemo_gym.base_responses_api_model.BaseResponsesAPIModelConfig()
```

**Bases:** [BaseRunServerInstanceConfig](/nemo/gym/nemo-gym/nemo_gym/config_types#nemo_gym-config_types-BaseRunServerInstanceConfig)

**`drop_custom_tools`** `bool`

---

**`drop_hosted_tools`** `bool`

---

**`token_id_capture_non_generating_requests`** `list[NonGeneratingRequest] = Field(default_factory=list)`

---

```python
class nemo_gym.base_responses_api_model.CaptureStore(
    root: str | pathlib.Path
)
```

Append-only, rollout-keyed JSONL sink for model exchanges.

**`_root`** `= Path(root)`

---

**`root`** `Path`

---

```python
nemo_gym.base_responses_api_model.CaptureStore.incomplete_path_for(
    rollout_id: str
) -> pathlib.Path
```

```python
nemo_gym.base_responses_api_model.CaptureStore.is_incomplete(
    rollout_id: str
) -> bool
```

```python
nemo_gym.base_responses_api_model.CaptureStore.mark_incomplete(
    rollout_id: str
) -> None
```

```python
nemo_gym.base_responses_api_model.CaptureStore.path_for(
    rollout_id: str
) -> pathlib.Path
```

```python
nemo_gym.base_responses_api_model.CaptureStore.read(
    rollout_id: str
) -> list[dict[str, typing.Any]]
```

```python
nemo_gym.base_responses_api_model.CaptureStore.read_available(
    rollout_id: str
) -> tuple[list[tuple[int, dict[str, typing.Any]]], int]
```

Read valid exchanges without letting one damaged line hide the rest.

```python
nemo_gym.base_responses_api_model.CaptureStore.record(
    rollout_id: str,
    exchange: dict[str, typing.Any]
) -> None
```

Append one exchange and fsync (durable across a killed box).

`flock` serializes appends to the same rollout across worker processes and threads while
allowing independent rollouts to write concurrently. This does blocking file IO + fsync,
so callers run it off the event loop (the capture middleware uses `asyncio.to_thread`).

```python
class nemo_gym.base_responses_api_model.ModelCallCaptureConfig()
```

**Bases:** `BaseModel`

Run-wide model-call capture settings from Gym's global config.

**`model_call_capture_dir`** `Optional[Path] = None`

---

**`observability_enabled`** `bool = False`

---

```python
nemo_gym.base_responses_api_model.ModelCallCaptureConfig.validate_capture_dir() -> nemo_gym.base_responses_api_model.ModelCallCaptureConfig
```

```python
class nemo_gym.base_responses_api_model.ModelCallRecord()
```

**Bases:** `BaseModel`

Observability record derived from one captured model-server exchange.

**`cache_creation_tokens`** `Optional[int] = None`

---

**`cache_hit`** `Optional[bool] = None`

---

**`cached_tokens`** `Optional[int] = None`

---

**`call_index`** `int`

---

**`client_session_id`** `Optional[str]`

---

**`completed_at`** `Optional[float] = None`

---

**`dialect`** `Optional[str] = None`

---

**`error_category`** `Optional[str] = None`

---

**`finish_reason`** `Optional[str] = None`

---

**`latency_total_ms`** `Optional[float] = None`

---

**`latency_ttft_ms`** `Optional[float] = None`

---

**`local_response_reason`** `Optional[str] = None`

---

**`model`** `Optional[str] = None`

---

**`model_call_id`** `Optional[str] = None`

---

**`model_ref`** `Optional[ModelServerRef] = None`

---

**`reasoning_content`** `Optional[str] = None`

---

**`request`** `Optional[dict[str, Any]] = None`

---

**`request_raw`** `Optional[str] = None`

---

**`response`** `Optional[dict[str, Any]] = None`

---

**`response_id`** `Optional[str] = None`

---

**`response_raw`** `Optional[str] = None`

---

**`response_source`** `Optional[Literal['upstream', 'local']] = None`

---

**`response_status`** `Optional[str] = None`

---

**`started_at`** `Optional[float] = None`

---

**`status_code`** `Optional[int] = None`

---

**`tokens_in`** `Optional[int] = None`

---

**`tokens_out`** `Optional[int] = None`

---

**`tokens_reasoning`** `Optional[int] = None`

---

**`tokens_total`** `Optional[int] = None`

---

**`tool_calls`** `list[dict[str, Any]] = Field(default_factory=list)`

---

**`upstream_attempted`** `Optional[bool] = None`

---

**`upstream_status_code`** `Optional[int] = None`

---

```python
class nemo_gym.base_responses_api_model.ModelExecutionOutcome
```

**Bases:** `typing.TypedDict`

**`error_category`** `NotRequired[str]`

---

**`local_response_reason`** `str | None`

---

**`response_source`** `Literal['upstream', 'local'] | None`

---

**`upstream_attempted`** `bool`

---

**`upstream_status_code`** `int | None`

---

```python
class nemo_gym.base_responses_api_model.SimpleResponsesAPIModel()
```

**Bases:** [BaseResponsesAPIModel](#nemo_gym-base_responses_api_model-BaseResponsesAPIModel), [SimpleServer](/nemo/gym/nemo-gym/nemo_gym/server_utils#nemo_gym-server_utils-SimpleServer)

**`non_generating_model_routes`** `frozenset[tuple[str, str]] = frozenset()`

---

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel._finalize_served_response(
    response: typing.Any
) -> None
```

async

Finalize capture after conversion to the response returned to the client.

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel._invoke_chat_completions(
    request: fastapi.Request,
    params: nemo_gym.openai_utils.NeMoGymChatCompletionCreateParamsNonStreaming
) -> nemo_gym.openai_utils.NeMoGymChatCompletion
```

async

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel._invoke_responses(
    request: fastapi.Request,
    params: nemo_gym.openai_utils.NeMoGymResponseCreateParamsNonStreaming
) -> nemo_gym.openai_utils.NeMoGymResponse
```

async

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel._stream_responses(
    request: fastapi.Request,
    params: nemo_gym.openai_utils.NeMoGymResponseCreateParamsNonStreaming,
    ns_map: nemo_gym.responses_streaming.NamespaceMap
) -> fastapi.responses.StreamingResponse
```

async

Serve the same converted response to capture and Responses SSE clients.

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel._stream_served_response(
    response: typing.Any,
    events: typing.Iterable[str | bytes]
) -> fastapi.responses.StreamingResponse
```

async

Serialize buffered SSE before committing externally staged capture.

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel.chat_completions(
    body: nemo_gym.openai_utils.NeMoGymChatCompletionCreateParamsNonStreaming = Body()
) -> nemo_gym.openai_utils.NeMoGymChatCompletion
```

async

abstract

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel.chat_completions_dispatch(
    request: fastapi.Request,
    body: dict = Body()
)
```

async

Default `/v1/chat/completions` entrypoint shared by every Gym model server.

A non-streaming request validates strictly against
`NeMoGymChatCompletionCreateParamsNonStreaming` and delegates to this server's own
`chat_completions()`, preserving the historical non-streaming behavior (including the
standard 422 shape). When the client sends `stream: true` (blackbox
Chat-Completions-over-SSE harnesses like the OpenClaw agent always do), the request is
sanitized onto that same strict model (drop `stream`/`stream_options`; see
`nemo_gym.chat_streaming`), validated identically, delegated to the same
`chat_completions()`, and the complete response is buffered and re-emitted as a
synthesized `chat.completion.chunk` SSE stream. Keepalive comments maintain the
connection while the backend computes; model output is still buffer-then-replay.

Only a genuine boolean `stream: true` takes the streaming path; any other value
(e.g. `"false"` or `1`) stays on the strict non-streaming path, which rejects the
malformed `stream` with the same 422 as before.

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel.messages(
    request: fastapi.Request,
    body: dict = Body()
)
```

async

Default Anthropic Messages \<-> Responses mapping shared by every Gym model server.

Translates the inbound Anthropic Messages request to the Responses API, delegates to this
server's own `responses()` (so it reuses whatever backend the server has), and maps the
result back to an Anthropic Messages response. When the client requested `stream: true`
(the Claude Code CLI always does), the complete response is re-emitted as a synthesized
Anthropic SSE event stream. Servers may override this for native Messages handling.

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel.responses(
    body: nemo_gym.openai_utils.NeMoGymResponseCreateParamsNonStreaming = Body()
) -> nemo_gym.openai_utils.NeMoGymResponse
```

async

abstract

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel.responses_dispatch(
    request: fastapi.Request,
    body: dict = Body()
)
```

async

Default `/v1/responses` entrypoint shared by every Gym model server.

A plain JSON request validates strictly against
`NeMoGymResponseCreateParamsNonStreaming` and delegates to this server's own
`responses()`, preserving the historical non-streaming behavior. When the client
requests `stream: true` (blackbox Responses-over-SSE harnesses like the Codex CLI
always do), the request is first sanitized from the streaming wire dialect (extra
bookkeeping fields, `namespace` tool specs — see `nemo_gym.responses_streaming`),
delegated to the same `responses()`, and the complete response is re-emitted as a
synthesized Responses SSE event stream. A `responses()` failure on this path is turned
into a terminal `response.failed` event rather than an HTTP 500 (bad-request validation
still fails eagerly, before the stream is committed).

```python
nemo_gym.base_responses_api_model.SimpleResponsesAPIModel.setup_webserver() -> fastapi.FastAPI
```

```python
class nemo_gym.base_responses_api_model._CaptureMiddleware(
    app: typing.Any,
    store: nemo_gym.base_responses_api_model.CaptureStore | None,
    model_server_name: str | None,
    token_store: typing.Any = None,
    configured_sink: typing.Any = None,
    lineage_store: nemo_gym.token_id_capture.protocols.LineageResolver | None = None,
    delta_records: bool = False,
    external_staging: bool = False,
    token_capture_enabled: bool = False,
    non_generating_requests: frozenset[tuple[str, str]] = frozenset()
)
```

Pure-ASGI per-rollout capture.

Always strips an optional `/ng-rollout/&lt;id&gt;` path prefix before routing (used as the capture
key) so the prefix is a stable routing feature independent of capture.
When `store` is set it buffers the request body and a copy of the response while forwarding both
downstream unchanged, so it composes with streaming (SSE) responses -- it never consumes or rewraps
the stream. SSE chunks are forwarded immediately except for the terminal event, which is released
after the capture is durable. Every chunk is also buffered for post-hoc reassembly, so a very long
stream is held in memory until it completes. When `store` is None it avoids buffering evaluation
records, while still persisting external capture failures before terminal events are forwarded.

**`_capture_ledger`** `CaptureLedger | None`

---

**`_lineage_store`** `LineageResolver | None = lineage_store`

---

```python
nemo_gym.base_responses_api_model._CaptureMiddleware.__call__(
    scope: dict[str, typing.Any],
    receive: typing.Any,
    send: typing.Any
) -> None
```

async

```python
nemo_gym.base_responses_api_model._as_arguments(
    arguments: typing.Any
) -> dict[str, typing.Any]
```

```python
nemo_gym.base_responses_api_model._cache_signal(
    usage: typing.Any
) -> tuple[typing.Optional[bool], typing.Optional[int]]
```

Cache hit/miss + cached-token count, from usage cache fields (OpenAI / Anthropic).

```python
nemo_gym.base_responses_api_model._classify_exception(
    exc: BaseException
) -> str
```

Normalized error\_category for an exception raised while calling the model.

```python
nemo_gym.base_responses_api_model._classify_status(
    status_code: int
) -> typing.Optional[str]
```

Normalized error\_category from an HTTP status (None when \< 400).

```python
nemo_gym.base_responses_api_model._consume_terminal_sse_event(
    buffer: bytearray,
    dialect: str
) -> typing.Optional[str]
```

```python
nemo_gym.base_responses_api_model._exception_http_details(
    exc: BaseException
) -> tuple[typing.Optional[int], bytes]
```

```python
nemo_gym.base_responses_api_model._fail_uncommitted_external_call(
    context: nemo_gym.token_id_capture.CaptureContext | None
) -> None
```

async

Record a failure when an admitted worker call returns without commit coordinates.

```python
nemo_gym.base_responses_api_model._headers_content_type(
    headers: list
) -> bytes
```

```python
nemo_gym.base_responses_api_model._orjson_dispatch_response(
    content: typing.Any
) -> typing.Any
```

Serialize a completed non-streaming model response with orjson.

FastAPI serializes a bare dictionary or Pydantic model with `jsonable_encoder`.
That function recursively visits every token ID and log probability before standard-library `json.dumps` runs.

Convert Pydantic models to JSON-compatible values before encoding the result.
Return the encoded bytes as a `Response` so FastAPI does not encode them again.
Preserve `Response` objects created by model-server overrides.

```python
nemo_gym.base_responses_api_model._parse_sse_events(
    raw: bytes
) -> list[dict[str, typing.Any]]
```

Parse an SSE byte stream into its JSON `data:` payloads (best-effort; non-JSON skipped).

```python
nemo_gym.base_responses_api_model._plain(
    value: typing.Any
) -> typing.Any
```

Reduce a request field to plain data so equal schemas serialize equally.

Handlers can expose tools as dictionaries or Pydantic models.
Normalization keeps an unchanged schema stable across handlers.

```python
nemo_gym.base_responses_api_model._preserve_capture_prefix_on_redirect(
    message: dict[str, typing.Any],
    capture_prefix: str,
    request_headers: list[tuple[bytes, bytes]]
) -> dict[str, typing.Any]
```

Keep rollout correlation on root-relative and same-origin redirects.

```python
nemo_gym.base_responses_api_model._reconstruct_anthropic_sse(
    events: list[dict[str, typing.Any]]
) -> typing.Optional[dict[str, typing.Any]]
```

Rebuild a complete Anthropic Messages response from its streamed events.

```python
nemo_gym.base_responses_api_model._reconstruct_chat_sse(
    events: list[dict[str, typing.Any]]
) -> typing.Optional[dict[str, typing.Any]]
```

Rebuild a Chat Completions response from streamed chunks.

```python
nemo_gym.base_responses_api_model._reconstruct_responses_sse(
    events: list[dict[str, typing.Any]]
) -> typing.Optional[dict[str, typing.Any]]
```

Rebuild a Responses API response: the terminal envelope carries the full response object.

```python
nemo_gym.base_responses_api_model._reconstruct_streamed_response(
    raw: bytes,
    dialect: str
) -> typing.Optional[dict[str, typing.Any]]
```

Best-effort: reassemble a final response object from a streamed (SSE) body, by dialect.

```python
nemo_gym.base_responses_api_model._record(
    store: nemo_gym.base_responses_api_model.CaptureStore,
    dialect: str,
    model_server_name: typing.Optional[str],
    request_bytes: bytes,
    client_session_id: typing.Optional[str] = None,
    rollout_id: str,
    model_call_id: str,
    started_at: float,
    completed_at: float,
    response_body: typing.Any,
    status_code: typing.Optional[int],
    error_category: typing.Optional[str],
    latency_ms: float,
    ttft_ms: typing.Optional[float] = None,
    response_raw: typing.Optional[str] = None,
    execution: typing.Optional[nemo_gym.base_responses_api_model.ModelExecutionOutcome] = None
) -> None
```

Append one exchange (success or failure). Best-effort: never raises.

```python
nemo_gym.base_responses_api_model._request_envelope(
    getter: typing.Any
) -> list[dict]
```

Return prompt-shaping request fields that are not turns.

Instructions and tools can be siblings of the message list.
The chat template renders both into the prompt.
Including them prevents prefix reuse across different request envelopes.

```python
nemo_gym.base_responses_api_model._request_messages(
    body: typing.Any
) -> list[dict]
```

Return the conversation carried by any supported dialect.

Lineage uses model-authored turns to identify the parent call.
Chat and Anthropic use `messages`.
Responses uses `input`.
The request envelope is prepended as a pseudo-turn because it also shapes the prompt.

```python
nemo_gym.base_responses_api_model._store_for_rollout(
    rollout_id: str,
    capture_dirs: list[pathlib.Path]
) -> typing.Optional[nemo_gym.base_responses_api_model.CaptureStore]
```

```python
nemo_gym.base_responses_api_model._token_count(
    value: typing.Any
) -> typing.Optional[int]
```

```python
nemo_gym.base_responses_api_model._tool_calls_and_reasoning(
    response: dict[str, typing.Any]
) -> tuple[list[dict[str, typing.Any]], typing.Optional[str]]
```

Structured tool calls (name, arguments, call\_id) and reasoning text, across all three shapes.

```python
nemo_gym.base_responses_api_model._unique_request_header(
    headers: list,
    name: bytes
) -> typing.Optional[str]
```

```python
nemo_gym.base_responses_api_model._usage_detail_token(
    usage: typing.Mapping[str, typing.Any],
    detail_groups: tuple[str, ...],
    field_names: tuple[str, ...]
) -> typing.Optional[int]
```

Return the first valid token count across equivalent provider detail shapes.

```python
nemo_gym.base_responses_api_model._validate_chat_params(
    body: dict
) -> nemo_gym.openai_utils.NeMoGymChatCompletionCreateParamsNonStreaming
```

Validate a /v1/chat/completions body dict, surfacing failures as FastAPI's standard 422.

`include_url=False` mirrors FastAPI's native body validation, which strips the
`errors.pydantic.dev` url from each detail entry, so the 422 body stays byte-for-byte
identical to the previous typed-`Body()` binding.

```python
nemo_gym.base_responses_api_model._validate_responses_params(
    body: dict
) -> nemo_gym.openai_utils.NeMoGymResponseCreateParamsNonStreaming
```

Validate a /v1/responses body dict, surfacing failures as FastAPI's standard 422.

```python
nemo_gym.base_responses_api_model._validate_rollout_id(
    rollout_id: str
) -> str
```

```python
nemo_gym.base_responses_api_model.aggregate_model_call_metrics(
    store: nemo_gym.base_responses_api_model.CaptureStore,
    rollout_id: str
) -> dict[str, typing.Any]
```

Aggregate model-call metrics for one rollout id.

```python
nemo_gym.base_responses_api_model.aggregate_model_call_records(
    calls: list[nemo_gym.base_responses_api_model.ModelCallRecord]
) -> dict[str, typing.Any]
```

Aggregate token and latency values from model-call records.

```python
nemo_gym.base_responses_api_model.build_model_call_record(
    exchange: dict[str, typing.Any],
    call_index: int
) -> nemo_gym.base_responses_api_model.ModelCallRecord
```

Map one captured exchange and its transport metadata into an observability record.

```python
nemo_gym.base_responses_api_model.clear_model_call_captures_for_rollouts(
    records: list[typing.Any],
    capture_dirs: list[pathlib.Path]
) -> None
```

Remove stale per-rollout capture files for these records before dispatch.

Capture files are keyed by a deterministic rollout id (task-rollout-attempt), so without this a
fresh run or a kill-shaped retry would append onto the previous attempt's capture for the same
id. The caller passes only rows about to be dispatched, after assigning any retry suffix.

```python
nemo_gym.base_responses_api_model.extract_token_stats(
    usage: typing.Any
) -> dict[str, typing.Optional[int]]
```

Normalize token totals across Responses, Chat Completions, and Anthropic Messages usage.

For native Anthropic `/v1/messages` with prompt caching, `input_tokens` is only the uncached
remainder, so cache-read + cache-creation tokens are folded into `tokens_in` to reflect the true
prompt size (and cache-creation is surfaced separately as `cache_creation_tokens`). OpenAI /
Responses usage already includes cached tokens in `input_tokens` / `prompt_tokens` (where
`cached_tokens` is a subset), so it is left untouched -- no double counting.

`tokens_in` is a prompt-*size* metric, not a cost proxy: providers price cache-read (\~0.1x) and
cache-creation (\~1.25x) differently from base input, so cost-accurate consumers should weight
`cached_tokens` and `cache_creation_tokens` separately rather than summing `tokens_in`.

```python
nemo_gym.base_responses_api_model.install_model_call_capture(
    app: typing.Any,
    config: nemo_gym.base_responses_api_model.ModelCallCaptureConfig,
    model_server_name: str | None = None,
    global_config_dict: typing.Any = None,
    num_workers: int | None = None,
    non_generating_requests: frozenset[tuple[str, str]] = frozenset()
) -> None
```

Install model-call capture middleware.

Always strip `/ng-rollout/&lt;id&gt;/...` before routing.
Evaluation capture records requests and responses for that path.
Non-terminal SSE chunks continue immediately.
The terminal event follows the durable evaluation write.
Training capture uses `/ng-rollout/&lt;id&gt;/training-token-capture/...`.
That path provides a request-scoped token sink.
The model server records token ids from its complete response.
Consumers access records through `TokenSource.freeze`.

```python
nemo_gym.base_responses_api_model.make_capture_store(
    config: nemo_gym.base_responses_api_model.ModelCallCaptureConfig
) -> typing.Optional[nemo_gym.base_responses_api_model.CaptureStore]
```

Build a CaptureStore when observability is enabled; otherwise None.

```python
nemo_gym.base_responses_api_model.merge_model_call_capture_into_record(
    record: dict[str, typing.Any],
    capture_dirs: list[pathlib.Path],
    include_payloads: bool = False
) -> dict[str, typing.Any]
```

Attach captured model-call observability data to a rollout record in place.

Keyed by the rollout id derived from the record's task/rollout/attempt indices, so the attached
shape is identical for every agent harness. Adds
`ng_model_call_capture = &#123;rollout_id, metrics, calls&#125;` where `calls` are derived observability
records. Raw request and response payloads remain in the capture store and are omitted from the
attachment unless `include_payloads` is true. Capture/read/join failures are attached as
`gaps`. The harness output and reward are not modified.

```python
nemo_gym.base_responses_api_model.model_call_capture_dirs_from_config(
    global_config_dict: typing.Any
) -> list[pathlib.Path]
```

Return the single run-wide capture directory when capture is enabled.

```python
nemo_gym.base_responses_api_model.observability_enabled_from_config(
    global_config_dict: typing.Any
) -> bool
```

Return the run-wide `observability_enabled` flag directly, without going through capture dirs.

```python
nemo_gym.base_responses_api_model.read_available_model_call_records(
    store: nemo_gym.base_responses_api_model.CaptureStore,
    rollout_id: str
) -> tuple[list[nemo_gym.base_responses_api_model.ModelCallRecord], int]
```

Read valid call records and count damaged records.

```python
nemo_gym.base_responses_api_model.read_model_call_records(
    store: nemo_gym.base_responses_api_model.CaptureStore,
    rollout_id: str
) -> list[nemo_gym.base_responses_api_model.ModelCallRecord]
```

Read captured exchanges in durable append order.

```python
nemo_gym.base_responses_api_model.start_model_execution(
    request: fastapi.Request,
    upstream_attempted: bool
) -> nemo_gym.base_responses_api_model.ModelExecutionOutcome
```

Keep adapter-owned execution facts separate from the served response and capture settings.

```python
nemo_gym.base_responses_api_model._ANTHROPIC_CONVERTER = AnthropicConverter()
```

```python
nemo_gym.base_responses_api_model._CHAT_KEEPALIVE_SECONDS = 15.0
```

```python
nemo_gym.base_responses_api_model._CLIENT_SESSION_HEADER = b'x-session-id'
```

```python
nemo_gym.base_responses_api_model._ENVELOPE_ROLE = '_ng_request_envelope'
```

```python
nemo_gym.base_responses_api_model._OBSERVED_PATHS = {'/v1/responses': 'responses', '/v1/chat/completions': 'chat', '/v1/messages': '...
```

```python
nemo_gym.base_responses_api_model._ROLLOUT_PATH_RE = re.compile(f'^/{re.escape(ROLLOUT_PATH_PREFIX)}/(?P<rollout_id>[^/]+)(?:/(?P<tok...
```

```python
nemo_gym.base_responses_api_model._SSE_KEEPALIVE = b': keep-alive\n\n'
```

```python
nemo_gym.base_responses_api_model._TERMINAL_SSE_LINES: dict[str, dict[bytes, str]] = {'responses': {b'event: response.completed': 'complete', b'event: response.incom...
```

```python
nemo_gym.base_responses_api_model.logger = logging.getLogger(__name__)
```