Understand Gateway Lifecycle Control
Built-in OpenClaw and Hermes images support two direct-container lifecycle topologies for recover and gateway restart.
The managed controller authenticates the host lifecycle action and prevents PID reuse from redirecting its signal.
It cannot prove process provenance against a malicious process running under the same sandbox UID, and it does not create gateway and agent UID isolation.
For recover and gateway restart, the managed controller acquires the expected-exit lock before it inspects the supervisor or gateway.
Lock acquisition, gateway termination, and replacement health share one recovery deadline.
If lock acquisition reaches that deadline, the controller returns SUPERVISOR_BUSY without publishing an expected-exit marker.
Hermes configuration remains mutable in both topologies. Before a restart, the supervisor validates the secret boundary and records one stable config snapshot in its transaction metadata. A direct MCP change becomes pending and supersedes stale host-managed intent. The supervisor marks that change as applied only after the replacement gateway passes its health checks. The host operation that owns a managed MCP transaction can report a registry mismatch after Hermes is healthy.
The direct root-entrypoint topology keeps the integrity metadata under root ownership while it seals a restart transaction. That metadata does not make the complete Hermes config a relaunch allowlist. The managed topology has no durable root-owned config anchor because the supervisor, gateway, and agent share one UID.
This compatibility path remains necessary while the OpenShell-managed topology owns a nonroot supervisor and shared gateway-agent UID. Remove it only after the minimum supported OpenShell provides a root-owned lifecycle supervisor or a gateway UID distinct from the agent, then migrate both built-in agents to that boundary.
Verify Recovery Health
For built-in OpenClaw and Hermes controllers, a successful recover or gateway restart response supplies the initial authenticated gateway-health proof.
After the settle window, NemoClaw sends one read-only authenticated probe through the same controller before it declares success.
The controller rechecks the managed child, listener, HTTP health, and required auxiliary processes from inside the gateway network namespace without restarting the gateway. A failed managed probe cannot be overridden by an outer-namespace HTTP response.
Custom agents that recover through an SSH script do not use this controller probe and continue to poll ordinary gateway health.
The nonroot Hermes supervisor continuously repairs the gateway, API relay, dashboard, dashboard relay, and gateway log stream. Four consecutive gateway health failures trigger recovery of the observed gateway child.
Five unexpected gateway exits or failed replacement candidates within 60 seconds stop relaunch for the current supervisor instance.
An authenticated host action authorizes one exit bound to the gateway process ID and kernel start identity while the root controller process remains live, so deliberate gateway restart and controller-driven replacement do not consume that crash budget.
Correct the reported process or health failure, then stop and start the sandbox to reset the supervisor. Rebuild only if the sandbox still cannot start.
The authorization records host intent for that exit; it does not claim that the host signal was the only possible cause of process termination in the shared-UID topology. After the in-sandbox processes are healthy, the host repairs only the host-side OpenShell forwards.
Fail Closed on Unsupported Topologies
The host selects the matching controller automatically for recover and gateway restart.
Ordinary openshell sandbox exec and manual in-sandbox relaunch are not fallback paths.
A current built-in image supports both the direct root-entrypoint and OpenShell-managed topologies.
An arbitrary nonroot entrypoint that does not match the managed OpenShell process shape fails closed with privileged control unavailable.
Kubernetes and other deployments without a matching direct container also fail closed with privileged control unavailable.
Older images without the matching supervisor or managed controller helper must be updated:
Related Topics
- Recover and Rebuild Sandboxes for recovery commands and rebuild fallback.
gateway restartorrecoverreportsprivileged control unavailablefor remediation.- Trusted Computing Base for the broader security boundary.