Recover and Rebuild Sandboxes
Use the lightest recovery operation that repairs the sandbox while preserving its supported state.
Restart a Stopped Sandbox Container
If status reports Phase: Error and confirms that the sandbox container exists but is stopped, restart the existing container:
This path preserves the sandbox workspace and repairs the agent runtime and host-side forwards after the container starts.
If the container is paused, follow the printed docker unpause guidance instead.
If the container is missing or OpenShell reports another terminal phase such as Failed, follow the printed rebuild --yes guidance so NemoClaw can recreate the sandbox from its recorded metadata.
Recover the Agent Runtime
Deep Agents sandboxes are terminal runtimes and do not expose an OpenClaw or Hermes in-sandbox gateway.
Use nemo-deepagents <sandbox-name> status, logs, connect, and rebuild for recovery.
If the terminal runtime reports degraded health, rebuild the sandbox instead of using recover or gateway restart.
Rebuild While Preserving State
If you changed the underlying Dockerfile, upgraded Deep Agents Code, enabled Tavily Search, or want to pick up a new base image without losing manifest-defined Deep Agents state, use rebuild instead of destroying and recreating.
Resolve Rebuild Preflight Stops
Before it backs up or deletes the existing sandbox, rebuild validates the recorded sandbox, gateway, policy, MCP, agent, and operation-lock state.
When one of these checks fails, NemoClaw prints Rebuild preflight failed, explains how to recover, and ends with Aborting rebuild.
At this boundary, the existing sandbox is unchanged and no sandbox data has been removed.
Use the recovery guidance that matches the reported check:
- Verify the sandbox name when its registry entry is missing.
- Follow the printed OpenShell gateway recovery steps when the gateway schema is incompatible.
- Repair the named pending baseline policy transition, then rerun
rebuild. - Resolve an incomplete MCP destroy transaction before retrying.
- Back up the sandbox state and recreate it with
nemo-deepagents onboardwhen the record contains multiple agents. Transactional multi-agent rebuild is not supported. - Wait for another onboarding or rebuild operation to finish before retrying. If verified stale-lock cleanup is still in progress, wait briefly and rerun the command. Do not delete the lock manually.
The rebuild command preserves manifest-defined Deep Agents state, regenerates config.toml, reconstructs managed MCP projection state, and reapplies registered policies while recreating the container.
Continue an Interrupted Replacement
Before rebuild deletes the existing sandbox, NemoClaw records a replacement journal in the onboarding session.
The journal binds the operation to the sandbox name, recorded OpenShell gateway, source identity, and replacement settings.
It stores fingerprints instead of credential values or raw OpenShell sandbox IDs.
If rebuild stops after recording the journal, rerun the command with the same replacement settings.
The rerun takes one of these actions:
- It continues deletion when the live sandbox still has the journaled source identity.
- It continues creation when the recorded OpenShell gateway explicitly reports the source sandbox as absent.
- It accepts an existing replacement only when its live identity and sandbox registry generation match the journal.
An accepted replacement is not deleted again.
The command reports Sandbox '<name>' already holds the replacement from the interrupted rebuild. and preserves the state backup path when one exists.
Pass --verbose to include the replacement identifier, OpenShell gateway, and journal phase in rebuild diagnostics.
After the sandbox registry proves the journaled replacement identity and generation, NemoClaw removes an obsolete source image that it owns.
It retains the image when the source is shared or the registered replacement reuses it.
If image removal fails, NemoClaw keeps the accepted replacement and tells you to run nemo-deepagents gc for cleanup.
NemoClaw fails closed when the selected gateway, replacement settings, durable source registry fields, or live source or target identity no longer matches the journal. The error names the sandbox and the mismatch that stopped recovery. Do not delete a same-name sandbox to bypass this check. Inspect the named OpenShell gateway and sandbox, correct the reported drift, and rerun the original command.
A same-name recreation started by nemo-deepagents onboard uses the same replacement journal.
If that recreation is interrupted after the Journaled replacement message, rerun the original onboarding command with the same target settings.
The active replacement can continue without adding --resume.
Use --resume for interrupted onboarding steps that occur before a replacement journal exists.
If an archive command preserves at least one state directory, NemoClaw keeps the usable entries and reports the manifest-defined paths that could not be archived.
If a manifest-declared state file fails, NemoClaw stops before deleting the original sandbox even when it preserved state directories, unless you explicitly pass --force.
If every state directory fails, NemoClaw stops before deleting the original sandbox even when it captured loose files, unless you explicitly pass --force.
rebuild --force can continue when no state directory was preserved or a manifest-declared state file failed.
NemoClaw restores any entries captured in the partial backup; if nothing usable was captured, it recreates the sandbox from recorded registry metadata without restoring prior sandbox state.
Use this recovery path only when losing the state that could not be backed up is acceptable.
When a sandbox with managed MCP servers cannot run a pre-mutation no-op, explicit --force uses its complete registry entries plus the exact live generated policies and provider identities to preserve MCP intent without scrubbing the unreachable in-sandbox adapter.
Every bridge entry must record the adapter for the sandbox’s recorded agent.
The registered policy must match the policy NemoClaw generates for that adapter, server name, endpoint URL, and resolved addresses.
NemoClaw rechecks that read-only snapshot immediately before deletion and stops if the target, registry, policy, provider, or recorded gateway changed.
NemoClaw sends the delete request and every deletion-confirmation lookup to the sandbox’s exact recorded gateway.
Across every rebuild path, NemoClaw does not attempt to stop the local NIM through the delete attempt, and cleanup is attempted on a best-effort basis only after deletion is positively confirmed.
After a nonzero delete, an explicit missing result converges as deleted.
A Ready or Running result triggers an attempt to restore prepared MCP state and any shields lockdown that rebuild temporarily opened.
NemoClaw reports any MCP or shields restoration failure and does not present the operation as a successful rollback.
Any partial or unreachable result remains ambiguous.
NemoClaw preserves the MCP ownership and rebuild-recovery records, does not attempt to stop NIM, skips the rebuild process’s immediate shields relock, and does not claim that the original sandbox is intact.
Inspect the live sandbox and gateway state before retrying recovery.
This recovery also stops for incomplete MCP adds or ambiguous ownership; an error after a successful no-op does not fall back to the host-side path.
When rebuild starts with shields up, NemoClaw opens a 30-minute shields-down window for backup and recreation. A detached auto-lock timer remains the recovery authority until NemoClaw commits a successful shields-up state, including when the host rebuild process exits unexpectedly.
Refer to nemo-deepagents <name> rebuild for flag details.
Use the Canonical Configuration Workflows
- Use Switch Inference Providers to change a model or provider.
Related Topics
- Create and Restore Snapshots for the state-preservation contract.
- Troubleshooting for
privileged control unavailable, stopped sandboxes, and failed rebuilds.