> For clean Markdown of any page, append .md to the page URL.
> For a complete documentation index, see https://docs.nvidia.com/openshell/llms.txt.
> For AI client integration (Claude Code, Cursor, etc.), connect to the MCP server at https://docs.nvidia.com/openshell/_mcp/server.

# Google Vertex AI

> Attach a Google Vertex AI provider and call native Vertex endpoints with gateway-refreshed credentials.

The `google-vertex-ai` provider gives selected sandboxes access to native
Google Vertex AI endpoints. OpenShell keeps refresh bootstrap material at the
gateway, rotates short-lived access tokens, and resolves token placeholders
only at endpoints authorized by the provider profile.

OpenShell does not choose a model or transform a request. The workload uses the
native Vertex endpoint and request format for its selected model.

## Prerequisites

* A GCP project with the Vertex AI API enabled.
* A service account with the Vertex AI User role and a downloaded JSON key for
  production, or gcloud Application Default Credentials for local development.
* Access to the selected model in the intended Vertex region.

## Create a Provider

### Service Account Key

Create the provider with the JSON key as gateway-only bootstrap material:

```shell
openshell provider create \
  --name vertex-prod \
  --type google-vertex-ai \
  --credential GOOGLE_SERVICE_ACCOUNT_KEY="$(cat /path/to/key.json)" \
  --config VERTEX_AI_PROJECT_ID=my-gcp-project \
  --config VERTEX_AI_REGION=us-central1
```

Configure gateway-managed refresh:

```shell
openshell provider refresh configure vertex-prod \
  --credential-key GOOGLE_VERTEX_AI_SERVICE_ACCOUNT_TOKEN \
  --strategy google-service-account-jwt \
  --material client_email="sa@my-gcp-project.iam.gserviceaccount.com" \
  --material private_key="$(jq -r .private_key /path/to/key.json)" \
  --secret-material-key private_key
```

The private key remains in the gateway credential store. Sandboxes receive
only an opaque placeholder for the short-lived access token.

### gcloud Application Default Credentials

For local development:

```shell
gcloud auth application-default login

openshell provider create \
  --name vertex-local \
  --type google-vertex-ai \
  --from-gcloud-adc \
  --config VERTEX_AI_PROJECT_ID=my-gcp-project \
  --config VERTEX_AI_REGION=us-central1
```

`--from-gcloud-adc` reads authorized-user ADC, configures an OAuth2 refresh
grant at the gateway, and immediately mints `GOOGLE_VERTEX_AI_TOKEN`. The ADC
file and refresh token do not enter the sandbox.

## Configuration Keys

| Key                    | Required | Default       | Description                                                   |
| ---------------------- | -------- | ------------- | ------------------------------------------------------------- |
| `VERTEX_AI_PROJECT_ID` | Yes      | —             | GCP project ID exposed as non-secret workload configuration.  |
| `VERTEX_AI_REGION`     | No       | `us-central1` | Vertex location exposed as non-secret workload configuration. |

When the provider is attached, OpenShell also projects standard project and
location aliases such as `GOOGLE_CLOUD_PROJECT`, `ANTHROPIC_VERTEX_PROJECT_ID`,
`CLOUD_ML_REGION`, and `VERTEX_LOCATION`.

## Attach the Provider

Attach it while creating a sandbox:

```shell
openshell sandbox create \
  --name vertex-agent \
  --provider vertex-local
```

Or attach it to an existing sandbox:

```shell
openshell sandbox provider attach vertex-agent vertex-local
```

Launch a new process after runtime attachment so it receives the provider
environment. Existing processes do not gain newly attached environment
variables.

## Call the Native Vertex API

Claude models use Vertex's publisher-model endpoint. Run a request from a new
sandbox process:

```shell
openshell sandbox exec vertex-agent -- sh -lc '
  token=${GOOGLE_VERTEX_AI_SERVICE_ACCOUNT_TOKEN:-$GOOGLE_VERTEX_AI_TOKEN}
  curl -X POST \
    -H "Authorization: Bearer $token" \
    -H "Content-Type: application/json" \
    -d '\''{
      "anthropic_version":"vertex-2023-10-16",
      "max_tokens":1024,
      "messages":[{"role":"user","content":"Hello"}]
    }'\'' \
    "https://${CLOUD_ML_REGION}-aiplatform.googleapis.com/v1/projects/${GOOGLE_CLOUD_PROJECT}/locations/${CLOUD_ML_REGION}/publishers/anthropic/models/claude-sonnet-4-6:rawPredict"
'
```

Use the model ID and location supported by your GCP project. For `global`, `us`,
or `eu`, use the corresponding Google-documented hostname instead of the
regional `<location>-aiplatform.googleapis.com` form.

Gemini and third-party models use their documented native or
OpenAI-compatible Vertex endpoints. Configure the model, URL, streaming mode,
and timeout in the client. OpenShell does not rewrite them.

## Verify and Troubleshoot

Inspect the attachment and effective policy:

```shell
openshell sandbox provider list vertex-agent
openshell policy get vertex-agent --full
openshell provider refresh status vertex-local
```

Common failures:

* A missing token variable usually means the process started before provider
  attachment. Launch a new process.
* `connection not allowed by policy` means the provider endpoint or caller
  binary is absent from the effective policy. A gateway global policy override
  suppresses provider-derived entries.
* `credential_endpoint_mismatch` means the request destination is outside the
  provider profile's endpoint binding.
* A Vertex 400 or 404 usually means the model, location, publisher path, or
  request body does not match the native API.
* A Vertex 401 or 403 can indicate an expired refresh grant or missing GCP IAM
  permission. Check `provider refresh status` and the Vertex AI User role.

Provider creation does not verify model access. The native request is the
end-to-end check.

## Migrate an Existing Vertex Route

An earlier managed route stored the provider and model separately and rewrote
requests for the workload. After upgrading, the provider and its refresh state
remain, but the route does not.

1. Attach the preserved Vertex provider to each intended sandbox.
2. Launch new workload processes.
3. Move the route's model and timeout into the client configuration.
4. Change the client to the native Vertex endpoint and request format.
5. Verify one non-streaming and one streaming native request before production
   rollout.

Do not attach the provider to every sandbox automatically. The old route was
workspace-global; the replacement intentionally grants access per sandbox.

## Next Steps

* [Provider-backed Inference](/sandboxes/inference-routing)
* [Profiles](/providers/profiles)
* [Providers](/sandboxes/manage-providers)
* [Customize Sandbox Policies](/sandboxes/policies)