Choose a Compatible Inference API

View as Markdown

Custom OpenAI-compatible endpoints use /v1/chat/completions at runtime by default. Choose the Responses API only when your endpoint implements the required streaming and tool-calling behavior.

Understand the Default Probe

During onboarding, NemoClaw probes /v1/responses first with tool-calling and streaming checks. It falls back to /v1/chat/completions when the Responses API does not provide the required behavior.

A successful Responses probe does not change the runtime API by itself. Without an explicit preference, the sandbox still uses /v1/chat/completions. This default avoids local backends that accept Responses requests but drop system prompts or tool definitions.

When a reasoning model returns only reasoning content before a final answer, NemoClaw retries the smoke request with a larger response budget. Route, configuration, and authentication failures still fail immediately.

Select the Responses API

Set NEMOCLAW_PREFERRED_API=openai-responses before onboarding.

NEMOCLAW_PREFERRED_API=openai-responses nemo-deepagents onboard

NemoClaw selects /v1/responses only when the validation response includes the required streaming events. If that probe fails, onboarding falls back to /v1/chat/completions automatically.

Select Chat Completions Only

Set NEMOCLAW_PREFERRED_API=openai-completions to skip the Responses probe and validate only /v1/chat/completions. This setting works in interactive and non-interactive onboarding.

NEMOCLAW_PREFERRED_API=openai-completions nemo-deepagents onboard
VariableValuesDefault
NEMOCLAW_PREFERRED_APIopenai-completions, openai-responsesUnset, which uses Chat Completions at runtime.

Understand the Deep Agents Runtime

NEMOCLAW_PREFERRED_API does not change the managed Deep Agents dcode runtime. Deep Agents sandboxes keep use_responses_api = false in /sandbox/.deepagents/config.toml and use Chat Completions through the OpenShell route.