Choose a Compatible Inference API
Custom OpenAI-compatible endpoints use /v1/chat/completions at runtime by default.
Choose the Responses API only when your endpoint implements the required streaming and tool-calling behavior.
Understand the Default Probe
During onboarding, NemoClaw probes /v1/responses first with tool-calling and streaming checks.
It falls back to /v1/chat/completions when the Responses API does not provide the required behavior.
A successful Responses probe does not change the runtime API by itself.
Without an explicit preference, the sandbox still uses /v1/chat/completions.
This default avoids local backends that accept Responses requests but drop system prompts or tool definitions.
When a reasoning model returns only reasoning content before a final answer, NemoClaw retries the smoke request with a larger response budget. Route, configuration, and authentication failures still fail immediately.
Select the Responses API
Set NEMOCLAW_PREFERRED_API=openai-responses before onboarding.
NemoClaw selects /v1/responses only when the validation response includes the required streaming events.
If that probe fails, onboarding falls back to /v1/chat/completions automatically.
Select Chat Completions Only
Set NEMOCLAW_PREFERRED_API=openai-completions to skip the Responses probe and validate only /v1/chat/completions.
This setting works in interactive and non-interactive onboarding.
Understand the Deep Agents Runtime
NEMOCLAW_PREFERRED_API does not change the managed Deep Agents dcode runtime.
Deep Agents sandboxes keep use_responses_api = false in /sandbox/.deepagents/config.toml and use Chat Completions through the OpenShell route.
Related Topics
- Set Up an OpenAI-Compatible Endpoint for endpoint configuration.
- Understand Provider Validation for validation behavior across providers.