About Inference Routing

View as Markdown

NemoClaw gives agents a managed inference path while OpenShell handles the selected upstream provider on the host. This design keeps provider credentials outside the sandbox.

Request Path

Most providers use the shared inference.local route inside the sandbox. OpenShell intercepts each request on the host and forwards it to the provider and model selected during onboarding.

Sandbox agent -> inference.local -> shared OpenShell route -> selected provider and model

NVIDIA Endpoints use a sandbox-attached OpenShell provider instead of the shared route. The agent addresses the NVIDIA API hostname, but OpenShell still intercepts the request and supplies the host-managed credential.

Sandbox agent -> integrate.api.nvidia.com -> sandbox-attached OpenShell provider -> NVIDIA

Host-side services such as Model Router and the OpenRouter runtime adapter remain behind the OpenShell route. The sandbox uses inference.local for those services instead of calling their host ports directly.

Credential Boundary

Provider credentials stay on the host and flow through the OpenShell provider system. The sandbox does not receive the raw upstream API key.

Local Ollama and local vLLM routes do not require the host OPENAI_API_KEY. NemoClaw uses provider-specific local tokens for those routes. Rebuilds of legacy local-inference sandboxes migrate away from stale OpenAI credential requirements. When a rebuild reuses an automatically bridged compatible-endpoint route without a host API key, NemoClaw reapplies the config-only bridge rewrite without reading or passing the credential stored in OpenShell.

Deep Agents Configuration

For Deep Agents, NemoClaw writes /sandbox/.deepagents/config.toml with a managed OpenAI-compatible provider, a scoped placeholder API key, and use_responses_api = false. Most providers use https://inference.local/v1. NVIDIA Endpoints instead uses the least-privilege OpenShell provider attached to the sandbox at https://integrate.api.nvidia.com/v1. The managed dcode runtime uses Chat Completions through OpenShell even when a compatible endpoint also supports the Responses API.