> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/nemoclaw/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/nemoclaw/_mcp/server.

# Configure Model Limits

> Context-window and output-token limits for a NemoClaw-managed sandbox.

Configure explicit model limits before onboarding so NemoClaw can bake them into the sandbox image.
Changing an explicit build-time model limit on an existing sandbox requires fresh recreation.

## Set the Hermes Context Window

Hermes accepts `NEMOCLAW_CONTEXT_WINDOW` as its model-limit override.

| Variable                  | Values                                    | Default                      |
| ------------------------- | ----------------------------------------- | ---------------------------- |
| `NEMOCLAW_CONTEXT_WINDOW` | Positive integer, at least `64000` tokens | Unset so Hermes auto-detects |

```bash
export NEMOCLAW_CONTEXT_WINDOW=65536
nemohermes onboard
```

When onboarding resolves a valid value, NemoClaw writes it as `model.context_length` in `/sandbox/.hermes/config.yaml`.
For non-Ollama endpoints, the field remains unset when no explicit or probed value is available so Hermes can auto-detect it from the endpoint.
During `inference set`, NemoClaw recomputes the context window for the target model.
It writes `model.context_length` when it resolves a value and omits the field when Hermes must use endpoint auto-discovery.
When NemoClaw starts Local Ollama on macOS or Linux, it requests at least `64000` tokens.
Fresh onboarding then verifies the loaded model's actual `context_length` through `/api/ps`.
Resumed onboarding and sandbox rebuilds warm the recorded Ollama model and repeat this verification before reusing its route.
When `NEMOCLAW_CONTEXT_WINDOW` is larger than `64000`, the Ollama runtime must provide at least that larger value.
When Ollama reports a smaller runtime value, NemoClaw queries `/api/show` for the selected model's native context window.
If the model's native context window is below the requirement, onboarding stops and tells you to select a model that meets the reported requirement.
If the model can meet the requirement, or NemoClaw cannot read its native context window, onboarding instead shows the required `OLLAMA_CONTEXT_LENGTH` value for restarting Ollama.
A missing or malformed runtime value also produces the daemon restart guidance.
Setting `NEMOCLAW_CONTEXT_WINDOW` does not raise the model's native context window or the Ollama daemon's runtime context, and it does not bypass this check.

## Use Detected Local Limits

When `NEMOCLAW_CONTEXT_WINDOW` is unset, NemoClaw can use a context length reported by the selected local server.
Local Ollama reports the loaded model's runtime context length.
Local vLLM and OpenAI-compatible endpoints can report `max_model_len` through `/v1/models`.

Set `NEMOCLAW_CONTEXT_WINDOW` when you need to override the detected value.

## Recreate an Existing Sandbox

Model limits are build-time settings.
Recreate the named sandbox after changing a supported value.

```bash
nemohermes onboard --fresh --name <sandbox-name> --recreate-sandbox
```

## Related Topics

* [Configure Inference Timeouts](configure-inference-timeouts) for request, validation, and readiness budgets.
* [Switch Models](switch-models) to change the selected model.