OpenSandbox Provider
The opensandbox provider creates sandboxes through the OpenSandbox SDK and server API. It is the Kubernetes-backed provider used by agentic software engineering environments such as mini_swe_agent_2.
Running at scale? Read OpenSandbox Best Practices for resource requests vs limits, timeout and retry guidance, and the failure signatures that are not bugs.
Setup
Obtain service access
To use Gym with OpenSandbox, connect to a running deployment and obtain its
service key. Set that key as OPENSANDBOX_API_KEY.
Choose an access route before running an agent:
- An existing organization-managed deployment: request access from the team operating that OpenSandbox service, using its approved internal access runbook.
- Your own deployment: follow the upstream
OpenSandbox server deployment guide
and server configuration reference.
The operator configures a non-empty
api_keyunder the server’s[server]section and securely supplies that value to clients asOPENSANDBOX_API_KEY. Deploy and secure the server before connecting Gym.
Configure the client
Install the sandbox extra in the runtime image or virtual environment that creates sandboxes:
For package installs, use:
Set connection values through the provider config. The shipped config reads these environment variables:
Sandbox, model, and image-registry credentials are separate. Keep secrets out of source control, issues, and resolved-config logs.
Check access before allocating sandboxes
From the Gym server’s host or container:
- Confirm the endpoint resolves and is reachable. The shipped
.svc.cluster.localaddress normally requires cluster network/DNS access. Use the operator’s TLS/CA configuration. - Use the public
/healthroute to check reachability, then the deployment’s read-only authenticated check to verify credentials. - Confirm image-registry and required download access, plus model connectivity
from the component making requests. Inside a sandbox,
localhostrefers to the sandbox itself. - Agree initial concurrency, lifetime, and cleanup within the deployment’s limits.
Resolve access, quota, or image-pull failures with the operator before retrying. After these checks, run an agent evaluation to verify the complete workflow.
Provider Config
Load the shipped OpenSandbox config
and select it with sandbox_provider: sandbox. This excerpt shows selected
defaults; load the full file using the launch command below.
Keep create.skip_health_check: false to avoid commands racing sandbox readiness.
Keep create.timeout_s above the task’s ready_timeout_s: the shipped mini-SWE
values are 1500 and 1200 seconds respectively. Adjust them together.
Run an agent with the provider config by passing it alongside the agent and model configs:
The provider constructor accepts four optional config sections:
SandboxSpec Provider Options
Set OpenSandbox-specific create options under SandboxSpec.provider_options, or under sandbox_spec.provider_options in an agent config:
Supported options are:
Unknown provider option keys raise ValueError; values with the wrong type raise TypeError. This catches config drift before a sandbox allocation is attempted.
Relevant SandboxSpec Fields
Resource Mapping
SandboxResources is translated into OpenSandbox resource quantities:
The provider also normalizes metadata values for backend labels by replacing unsupported characters and truncating values to the provider limit.
Job Attribution
The provider merges nemo-gym.nvidia.com/team, nemo-gym.nvidia.com/user,
nemo-gym.nvidia.com/workload, and nemo-gym.nvidia.com/run keys into each sandbox’s
metadata when they can be resolved. OpenSandbox propagates sandbox metadata as Kubernetes
labels on the sandbox resources, so running sandboxes are attributable both through the
OpenSandbox list API metadata filter and directly at the cluster level:
Each field resolves in order: the provider’s attribution config, then NEMO_GYM_TEAM /
NEMO_GYM_USER / NEMO_GYM_WORKLOAD environment variables, then Slurm job environment
variables (SLURM_JOB_ACCOUNT / SLURM_JOB_USER / SLURM_JOB_NAME). user additionally
falls back to the OS login name (root is ignored — containers run as root by default, so
it would attribute the image, not a person), and workload to NEMO_GYM_CONFIG_PATH, the
server instance name the gym CLI sets on every server process it spawns. Fields that cannot
be resolved are omitted. Explicit sandbox_spec.metadata and default_metadata keys always
take precedence over attribution keys.
run identifies one launch of the creating process: NEMO_GYM_RUN_ID if set, else an id
generated per process. The resolved attribution is logged once at the first sandbox create —
note the run value from the logs to list or clean up exactly that run’s sandboxes (for
example after killing a run partway through), which team / user / workload cannot do
when the same user launches the same workload twice:
Attribution in Kubernetes deployments
When the gym servers themselves run in Kubernetes (for example a driver pod launching an
eval), none of the automatic sources resolve: there are no Slurm variables, and the
container’s OS login is typically root, which is ignored. Set the NEMO_GYM_* variables
on the pod spec — either directly or via the downward API from labels the pod already
carries:
Lifecycle
The provider creates one OpenSandbox sandbox per Gym sandbox:
If create.skip_health_check is enabled, the provider reconnects to the sandbox before running the NeMo Gym readiness probe so follow-up operations use a fresh SDK handle.
File Transfer
The provider uses the OpenSandbox SDK file API for uploads and downloads:
upload()reads the local file and writes bytes to the target sandbox path.download()reads bytes from the sandbox path and writes them to the local target path.
Startup files from SandboxSpec.files use the same provider file path semantics before the first command runs.
User and Runtime Notes
The neutral user argument to exec() maps onto OpenSandbox command options:
Command failures return SandboxExecResult with the command’s exit code. If OpenSandbox reports an execution error without an exit code, the provider returns code 125 with error_type="sandbox".
The provider’s create.image_pull_policy defaults to IfNotPresent. Valid values are Always, IfNotPresent, and Never. The resolved policy is written into OpenSandbox create extensions as both imagePullPolicy and opensandbox.extensions.image-pull-policy unless those keys are already present in provider_options.extensions.
Create retries handle transient allocation, connection, image pull, and server-side errors. Command retries are controlled separately by operations.command_retries.
execd, the exec daemon inside each sandbox, holds every command response open for EXECD_API_GRACE_SHUTDOWN (default 1s) after the command’s last event, which is fixed latency on each command and on the readiness probe. Set it in the sandbox spec’s env to shorten it; the shipped SWE-bench and OpenCode configs use 50ms, and the provider’s operations.background_poll_initial_s (0.5) is sized to that.
connection.tls_verify (default false) controls certificate verification on every connection the provider opens, both the SDK transport and PTY WebSockets. The default skips verification so that https:// endpoints with a certificate the client cannot verify, such as a validation cluster behind a self-signed load balancer, work without extra setup; set tls_verify: true for endpoints with a verifiable certificate. For https endpoints, either set connection.protocol: https or include the scheme in the domain (connection.domain: https://sandbox.example, or OPENSANDBOX_DOMAIN=https://sandbox.example): a scheme in the domain sets the protocol and takes precedence over connection.protocol. The cleanup_sandboxes.py helper mirrors the setting with --tls-verify and reads tls_verify from --connection-config.
Setting connection.keepalive_expiry_s: null preserves the configured tls_verify behavior. It uses the SDK’s default transport only when tls_verify: true and disable_connection_pooling: false; otherwise, the provider supplies a transport to apply those settings.