NemoClaw Architecture Overview

View as Markdown

This page explains how NemoClaw runs supported agent runtimes inside OpenShell sandboxes. It covers the host CLI, OpenShell gateway, agent integration layer, lifecycle state, managed Model Context Protocol (MCP) servers and other integrations, and protection layers.

NemoClaw does not replace OpenShell or the selected agent runtime. It packages them as a repeatable setup with a versioned blueprint, agent-specific configuration, managed inference, network policy, and lifecycle operations.

High-Level Flow

NemoClaw keeps operator control on the host while OpenShell enforces the sandbox boundary. The OpenShell gateway coordinates sandbox lifecycle, credentials, network policy, inference routes, and approved integration traffic.

The diagram has the following components:

ComponentRole in the flow
Users and operatorsInstall and operate NemoClaw from the host, then interact through the selected agent interface.
NemoClaw host CLICollects configuration, runs readiness checks and onboarding, resolves the blueprint, and operates managed resources.
OpenShell gatewayCoordinates sandbox lifecycle, credentials, networking, policy enforcement, inference routing, and approved integration egress.
OpenShell sandboxRuns the selected agent runtime with its NemoClaw integration layer, configuration, and supporting tools.
Agent interfaceProvides the interaction path exposed by the selected agent runtime.
Inference providersReceive managed inference requests through the OpenShell gateway.
Approved integrationsReceive policy-approved requests to MCP servers, package indexes, and other configured services.
Managed state and artifactsPreserve non-policy registry records, workspace files, logs, and manifest-declared snapshot content; OpenShell alone stores sandbox policy.

For repository layout, file paths, and deeper diagrams, refer to Architecture.

Design Principles

NemoClaw follows these architecture principles.

Versioned blueprint : The blueprint runner resolves a versioned blueprint and verifies its digest before it changes managed resources.

Host credential custody : OpenShell stores inference provider credentials and managed MCP bearer values outside the sandbox and replaces placeholders at approved request boundaries.

Agent-specific integration : Each supported agent runtime receives the configuration, wrappers, plugin, or adapter required for its documented workflow.

Resumable lifecycle : NemoClaw records lifecycle progress and reconciles managed resources after supported interruptions or partial operations.

Manifest-declared state : Rebuild, snapshot, and restore operations preserve only the state declared for the selected agent runtime. Each agent manifest and operation defines which credential-bearing files to exclude.

Host-configured messaging credentials also use OpenShell credential delivery. Some messaging integrations, such as QR-paired WhatsApp, retain explicitly declared session credentials inside the sandbox so supported lifecycle operations can preserve them.

CLI, Integration Layer, and Blueprint

NemoClaw separates host orchestration, agent-specific behavior, and sandbox definition.

  • The host CLI runs readiness checks and onboarding, validates provider choices, records lifecycle state, and operates OpenShell resources.

  • The Hermes integration layer writes runtime configuration under /sandbox/.hermes, including config.yaml, environment files, and supported messaging-channel settings.

  • The blueprint is a versioned YAML package with the sandbox image, agent manifest, network policy, inference profile, and supporting assets. The runner resolves and verifies the blueprint before applying it through OpenShell.

This separation keeps host orchestration, agent-specific assets, and the sandbox definition at explicit lifecycle boundaries.

Inspect an External OpenShell Gateway

NemoClaw packages the experimental nemoclaw-blueprint-runner command for infrastructure that already owns an external OpenShell gateway. This path supports target planning and one credential-free public health request for OpenShell 0.0.106. It does not manage the gateway or establish support for another OpenShell release.

Create a blueprint directory with a blueprint.yaml file that contains one external target. The target must use a bare HTTPS origin, the exact OpenShell release, an absolute CA bundle path, and an absolute authentication file path. The CA bundle must be a nonempty regular file, must not be a symbolic link, and must be no larger than 1 MiB. The CA bundle must contain only PEM CA certificates. Both operations read the CA bundle. The plan requires a nonempty regular authentication file no larger than 1 MiB and inspects only its metadata. The status command does not access the authentication file. Neither operation creates or removes either administrator-owned file.

1version: 1.0.0
2min_openshell_version: 0.0.106
3max_openshell_version: 0.0.106
4openshell_target:
5 endpoint: https://openshell.example.test:8443
6 workspace: default
7 expected_release: 0.0.106
8 lifecycle: external
9 trust:
10 ca_file: /var/run/openshell-target/ca.pem
11 authentication:
12 credential_file: /var/run/openshell-target/authentication

Run the target-only plan when you need to validate the complete target configuration and print a sanitized plan. Planning requires the referenced authentication file, validates local input, fingerprints the CA bundle, and does not connect.

$NEMOCLAW_BLUEPRINT_PATH=/absolute/path/to/blueprint nemoclaw-blueprint-runner plan

The status command sends external traffic to the configured gateway. The OpenShell SDK uses platform DNS without IP address pinning and verifies the server certificate and hostname against the supplied CA bundle. You must control the target hostname and its DNS resolution.

Run the external status command directly when you only need public health. Status requires the absolute authentication file reference but does not access the file. It makes one public health request through the official OpenShell SDK.

$NEMOCLAW_BLUEPRINT_PATH=/absolute/path/to/blueprint nemoclaw-blueprint-runner status --external-target

A successful status result reports healthy, release 0.0.106, and compatible in JSON. The command does not access the authentication file or authenticate. The command supplies the validated CA bundle to the SDK and uses normal TLS hostname verification.

This experimental path does not create, update, delete, or list gateways, workspaces, sandboxes, credentials, or policies. It does not establish workspace readiness or support machine authentication or Kubernetes. A failed status request makes no remote change, so correct the endpoint, CA bundle or its file metadata, gateway health, or release mismatch before you retry. If status reports that the approved OpenShell SDK is unavailable, restore a NemoClaw installation that includes @nvidia/openshell-sdk 0.0.106 before retrying.

Readiness and Sandbox Creation

Run nemohermes host probe when you need a read-only system readiness report before onboarding. The report combines host and gateway observations, capabilities, qualifications, findings, evidence, and CLI provenance without changing system state. Onboarding consumes the same stable host and gateway entities and applies its explicit admission policy. It revalidates live facts after permitted preparation and when a saved onboarding session resumes.

When you run nemohermes onboard, the host CLI and blueprint runner complete these operations:

  1. NemoClaw resolves gateway lifecycle authority and rejects blocking system readiness results before managed resource effects. A container-backed WSL GPU proof can run only after this admission check; explicit CPU-only intent skips it.
  2. NemoClaw resolves the blueprint, checks version compatibility, and verifies the digest.
  3. Onboarding validates the selected inference provider, credentials, agent settings, and platform requirements.
  4. The runner determines which gateway, provider, policy, sandbox, and integration resources to create or update.
  5. NemoClaw records progress so a supported interruption can resume or report a specific recovery action.

Before the blueprint runner writes a temporary policy update, it parses the merged live policy and refuses a literal credential value. Replace literal credentials with supported OpenShell credential bindings or resolver placeholders, then retry. The refusal happens before the policy file or OpenShell policy mutation is created.

After the sandbox starts, the selected agent uses its managed configuration and the controls supported by the host.

Lifecycle and State

NemoClaw operates the sandbox and its manifest-declared state through host-side commands.

OperationResult
Inspecthost probe, status, and logs report system, sandbox, agent-runtime, inference, and recovery information without replacing the sandbox.
ConfigureInference, policy, managed MCP, and supported agent-runtime integration commands update the applicable managed resources.
RebuildRecreates the sandbox from the recorded configuration and restores supported agent state through a recorded transaction.
RecoverRepairs a stopped or degraded agent runtime and its sandbox-scoped forwards when the recorded identities still match.
Snapshot and restoreCaptures manifest-declared state with the agent- and operation-specific credential exclusions, then applies that state to an eligible sandbox.
Destroy and uninstallRemoves the selected sandbox or host installation according to the command scope and preservation choices.

Refer to Recover and Rebuild Sandboxes and Create and Restore Snapshots for lifecycle details.

Inference Routing

Managed agent runtimes send model requests to inference.local instead of an upstream endpoint. During onboarding, NemoClaw validates the selected provider and model, configures the OpenShell inference route, and writes the matching model reference into the managed agent configuration. OpenShell keeps the provider credential outside the sandbox and sends approved requests to the upstream endpoint. When you select the Model Router provider, inference.local routes to a host-side router that chooses from the configured NVIDIA model pool for each request.

For Hermes, nemohermes inference set updates /sandbox/.hermes/config.yaml at runtime without rebuilding the sandbox.

Managed Integrations

NemoClaw connects supported external services through OpenShell providers, network policy, and agent-specific adapters.

Managed MCP supports authenticated HTTPS Streamable HTTP MCP servers for OpenClaw, Hermes, and Deep Agents Code. NemoClaw stores the credential name and ownership metadata, while OpenShell stores the raw value outside the sandbox. The agent adapter receives a credential placeholder that OpenShell replaces only at the approved egress boundary.

Messaging channels use agent-specific channel manifests, credential delivery, network policy, and lifecycle commands. Some experimental webhook channels also require a route-restricted host-side public endpoint. Refer to Choose Messaging Channels for agent and channel status.

Refer to About Managed MCP Servers for the managed MCP security and lifecycle design.

Protection Layers

The sandbox starts with a baseline policy that controls network egress, filesystem access, process privileges, and inference routing.

LayerWhat it protectsWhen it applies
NetworkBlocks unauthorized outbound connections.Hot-reloadable at runtime.
FilesystemRestricts system paths to read-only; /sandbox and /tmp are writable.Locked at sandbox creation.
ProcessBlocks privilege escalation and dangerous syscalls.Locked at sandbox creation.
InferenceReroutes model API calls to controlled backends.Hot-reloadable at runtime.

When the agent tries to reach an unapproved host, OpenShell blocks the request and surfaces it in the terminal user interface (TUI) for operator approval. Approved endpoints persist within the current sandbox instance but are not saved to the baseline policy file. NemoClaw’s runtime context tells supported agents to try allowed network and filesystem actions first, then report whether policy denial, DNS, timeout, TLS, or filesystem access caused a failure.

Host and platform limitations can change how individual controls apply. Refer to Platform Support and Security Best Practices before you treat a control as an environment-wide guarantee.

Next Steps