Deploy Agents
Deploy a registered agent as a running service and invoke it through the Agents gateway. An agent can run in one of three modes — as a local subprocess (the default), or as a durable container on Docker or Kubernetes.
Resource names for agents and deployments must contain only letters (a-z,
A-Z), digits (0-9), underscores, hyphens, and dots. For example:
calculator-agent, my-agent, react-agent.
CLI
Python SDK
Deployment Modes
nemo agents deploy --mode <mode> selects the runtime backend for a
deployment. The default is subprocess; docker and k8s run the agent as a
durable container through the deployments plugin.
In every mode the agent is reached the same way — through the Agents gateway, which resolves the agent’s active deployment and proxies the request, so clients do not need to know which mode the agent runs in.
Agent Environments
An agent config declares what an agent needs (MCP servers, environment variables, and the names of the credentials it reads) without pinning any of it to a place to run. An agent environment supplies those specifics for a given context. This allows one agent to be reused across different contexts. For instance, a single agent can run with dev credentials and full tool access in one environment, and with more restricted credentials, read-only tools, and a larger compute box in another environment.
An environment is built from three resources:
Secrets are referenced by name (workspace/secret-name) and resolved from the
platform Secrets service
at run time, so credential values never live in the agent config or the
environment spec.
Example
Define the Agent
The agent config declares a github MCP server and the env var it reads its
credential from. It holds no secret and names no environment. See
Agent Definition for the full
agent.yaml reference.
Register it:
CLI
Python SDK
Create an Environment
The environment fulfills the agent’s github server with a stored secret and a
tool scope, and provides an inline compute size. Create the secret first, then
the environment-spec, then the environment that references it.
CLI
Python SDK
Reference the Environment in an AgentDeployment
Deploy the agent under the environment. Its spec is merged into the agent config, and its secret refs and compute are snapshotted onto the deployment. Where both set the same field, the environment-spec’s value wins.
CLI
Python SDK
Subprocess Mode (Default)
The simplest path: the platform launches a FastAPI server for the
agent on its own host, assigns a port, watches its health, and tears it down
on nemo agents undeploy. No image or executor configuration is required.
CLI
Python SDK
Container Modes (Docker / Kubernetes)
Container modes give an agent a durable deployment that survives a platform restart. Instead of a local process, the platform compiles the agent into a generic deployment and hands it to the deployments plugin, which runs it on the configured executor (Docker or Kubernetes) and projects the running container’s address back onto the agent deployment. The Agents gateway then routes to that projected address.
Prerequisites
1. A container image for the agent. Container modes run a packaged agent
runtime with the selected harness adapters and the agent’s dependencies. Build
one with nemo agents package; the command detects nemo-agents-spec-v1 and
selects the Platform agent image pipeline automatically. Image building
requires the container extra. Install it with:
The examples below use the calculator agent that ships with the source
checkout. Run them from the repository root. Its config is located at
plugins/nemo-agents/examples/nemo-agent-config/calculator-agent/agent.yaml.
CLI
--mode docker needs an image the platform’s Docker daemon can run.
--mode k8s needs an image the cluster nodes can pull (a registry image, or an
image pre-loaded onto the nodes) — the k8s backend does not use image pull
secrets. Pass the image with --image, or set
agents.deployments.default_image in the platform configuration.
For hand-built images that already contain an agent server, pass
--use-image-entrypoint so the deployment preserves the image ENTRYPOINT/CMD
instead of injecting the platform-packaged agent server command.
The image must bind 0.0.0.0:$PORT, serve /health, and read config from
AGENT_CONFIG_PATH for Fabric or NAT_CONFIG_PATH for NAT.
2. A configured deployments executor. The platform operator defines named executors in the platform configuration and points the agents plugin at them. A minimal Docker + Kubernetes configuration:
The Kubernetes backend requires the kubernetes Python client in the platform
image and a ServiceAccount with permission to manage Deployments, Services,
ConfigMaps, and Pods in the target namespace. In the packaged Helm chart these
run in the core controller, whose Role already grants those permissions.
Deploy on Docker
CLI
Python SDK
Deploy on Kubernetes
Deployment is identical apart from --mode k8s. The deployments plugin creates
a Kubernetes Deployment and a ClusterIP Service, and projects the Service’s
in-cluster DNS address (<service>.<namespace>.svc.cluster.local:<port>) onto
the agent deployment. Because the Agents gateway runs in-cluster, it routes to
that address directly.
CLI
Python SDK
Model Access from a Deployed Agent
Regardless of mode, model traffic from inside the agent routes back through the
Inference Gateway. The platform injects
the gateway URL when it deploys the agent, and the gateway resolves model
entity names to upstream providers and supplies their credentials. Two
conventions apply to agent.yaml:
- Set
models.default.modelto the Inference Gateway entity name. The models controller creates these names by replacing slashes and dots with hyphens (nvidia/nemotron-3-nano-30b-a3bbecomesnvidia-nemotron-3-nano-30b-a3b). - Leave
base_urlunset for a Platform-routed model. Whenproviderisnvidia,openai, oropenai-compatible, the deployment supplies the Inference Gateway URL.api_key_envnames the environment variable expected by the selected harness; it does not contain a credential.
The calculator agent uses:
To make an external model available to the agent, register a provider first — see Deploy Models for NVIDIA Build, OpenAI, and Anthropic examples.
A deployed agent needs to reach the platform from inside its container. The
deployment handles this automatically for both Docker and Kubernetes, so you
normally don’t need to configure anything. If an agent can’t reach the platform,
set agents.deployments.gateway_url_override to a URL that is reachable from
inside the container.
Docker mode on Linux
On Linux, host.docker.internal doesn’t resolve inside containers, so agent
invokes can fail with openai.APIConnectionError. Point the deployment at the
Docker bridge address (172.17.0.1 by default) in config.yaml, then start the
platform bound to all interfaces:
If your Docker bridge uses a non-default subnet, substitute its gateway address
(docker network inspect bridge --format '{{ (index .IPAM.Config 0).Gateway }}').
Inspect a Deployment
CLI
Python SDK
For a container-mode deployment, the deployment reports deployment_mode
(docker or k8s), a status of running once ready, and an endpoints
list carrying the container’s routable address. Subprocess deployments carry a
loopback endpoint instead. The Agents gateway uses whichever the deployment’s
mode provides, so invocation is identical across modes.