Set Up an OpenAI-Compatible Endpoint
Set Up an OpenAI-Compatible Endpoint
Use the custom OpenAI-compatible provider for servers that implement /v1/chat/completions or a compatible /v1/responses API.
Examples include vLLM, TensorRT-LLM, llama.cpp, LocalAI, and other compatible servers.
The agent connects to inference.local inside the sandbox.
OpenShell forwards that traffic to the endpoint configured during onboarding.
Start the Server
Start the compatible server before onboarding.
The following example starts vLLM on port 8000.
Port 8000 is the default vLLM host-gateway port.
Set NEMOCLAW_VLLM_PORT before onboarding to select a different port.
For the no-authentication path on bundled host-gateway ports, bind the backend to loopback only. NemoClaw places a token-protected proxy in front of the endpoint. At startup, NemoClaw normally inspects the backend listeners. NemoClaw rejects an observed non-loopback listener because it would bypass the proxy token check. If listener inspection is unavailable, the proxy starts without verifying the backend bind. Onboarding does not display this degraded result. Before you continue, inspect the configured backend port:
Continue only when every listener address is in 127.0.0.0/8, is ::1, or is an IPv4-mapped address in ::ffff:127.0.0.0/8.
Stop the server and correct its bind configuration if the output contains *, 0.0.0.0, ::, or another address.
Endpoints configured with COMPATIBLE_API_KEY use a different authenticated path and do not use this loopback-only proxy requirement.
Run Onboarding
Start the onboard wizard.
Select Other OpenAI-compatible endpoint. Enter the server base URL and the model ID reported by the server.
Use a host-routable URL such as http://localhost:8000/v1 when you want onboarding to verify the API, tool-calling, and streaming paths before gateway registration.
To qualify for automatic rewriting, an HTTP endpoint URL must use the loopback host localhost, 127.0.0.1, or [::1].
Automatic rewriting is limited to the port selected by NEMOCLAW_VLLM_PORT (8000 by default) and ports 11434 and 11435.
NemoClaw validates the entered URL from the host and registers the OpenShell gateway route through host.openshell.internal:<port> for sandbox traffic.
Sandbox inference requests continue to use the base inference.local policy, so the managed compatible-endpoint route does not require adding the local-inference preset.
For authenticated endpoints, NemoClaw leaves URLs without an explicit port, URLs on :80 or another privileged port, and URLs on unsupported ports unchanged.
Those authenticated URLs require a separately compatible runtime topology and network policy.
This rewrite depends on an OpenShell topology that resolves host.openshell.internal inside the sandbox; if that bridge is unavailable, onboarding can still validate the host URL, but nemoclaw <name> status is the authoritative runtime check.
For no-authentication endpoints on the port selected by NEMOCLAW_VLLM_PORT (8000 by default) or port 11434, keep the server bound to loopback.
NemoClaw’s token-protected proxy makes the loopback service reachable from the sandbox without exposing the backend on other host interfaces.
Port 11435 is reserved for the proxy and is unavailable for new no-authentication endpoints.
Recovery can retain an existing no-authentication route on 11435 only when the proxy has moved and no other protected service owns that port.
The shared proxy stays pinned to its existing backend; reuse that endpoint.
Moving or removing a sandbox does not release the host-global binding.
To select a different no-authentication backend, back up the sandboxes and remove the final NemoClaw gateway with the uninstaller before reinstalling.
If you manually enter a sandbox-internal alias such as http://host.openshell.internal:8000/v1, host-side endpoint probing is skipped during onboarding.
Use a host-routable endpoint such as localhost when you need onboarding to verify the API, tool-calling, and streaming paths before gateway registration.
Otherwise, verify the runtime route after onboarding with nemoclaw <name> status and a short agent request.
For an HTTP URL using the host localhost, 127.0.0.1, or [::1] and the port selected by NEMOCLAW_VLLM_PORT (8000 by default) or port 11434, the API key prompt says that pressing Enter selects no authentication.
Other URLs still require COMPATIBLE_API_KEY.
Refer to Choose a Compatible Inference API for the probe order and runtime API selection.
Serve a Raw Model File
Start a compatible server for a raw model file instead of passing the file path to NemoClaw.
The Ollama provider accepts Ollama model tags and does not accept a raw .gguf path.
The following example starts llama-server with a GGUF model.
During onboarding, select Other OpenAI-compatible endpoint.
Enter the server base URL and the model ID returned by /v1/models.
Use the model ID, not the raw file path.
For the example above, the server commonly reports NVIDIA-Nemotron3-Nano-4B-Q4_K_M.gguf as its model ID.
Run Non-Interactive Onboarding
Start the endpoint before running the non-interactive command because onboarding validates the server.
Set NEMOCLAW_REASONING=true when the endpoint serves a reasoning-only model.
Reasoning mode validates only /v1/chat/completions and does not verify tool calling or streaming.
Enable it only when the endpoint supports the capabilities your agent requires.
For the raw model example, use the ID returned by /v1/models.
A loopback endpoint onboarded in no-auth mode keeps that binding after onboarding.
nemoclaw <sandbox-name> inference set accepts the recorded endpoint URL, or no --endpoint-url at all, and verifies the new route from inside the sandbox because the gateway reaches the endpoint through a local proxy.
Omit --credential-env: the recorded no-auth binding is the only value this route accepts, and only onboarding can rebuild that binding.
For OpenClaw, NEMOCLAW_REASONING_EFFORT accepts low, medium, high, or default.
A low, medium, or high value writes params.extra_body.reasoning_effort when the selected API is openai-completions.
Another API family omits the field.
An unset value or default leaves the endpoint’s own default in place.
NemoClaw rejects an invalid value or a provider/API mismatch before changing provider, sandbox, policy, or registry state.
After inference set changes the effort, an ordinary sandbox restart preserves the persisted value instead of restoring the image’s original onboarding value.
Private and reserved addresses are blocked by default. To use an inference gateway on a trusted corporate network, list only its host and keep the endpoint URL on that host:
NemoClaw still resolves the host before probing and pins outbound validation to the complete canonical address set. A trusted host can return both public and supported private addresses. NemoClaw pins every canonical answer. If any answer is a disallowed private, reserved, or special-purpose address, validation rejects the endpoint instead of discarding that answer. Among private answers, NemoClaw admits only RFC1918, carrier-grade network address translation (CGNAT), and IPv6 unique local address (ULA) destinations. Link-local metadata and other reserved ranges remain blocked. An unlisted private host, a hostname suffix match, or a DNS failure also remains blocked.
Related Topics
- Choose a Compatible Inference API to select Chat Completions or Responses.
- Meet Custom Endpoint Security Requirements before saving a public custom endpoint.
- Verify the Inference Route after setup.