google-agents-cli-deploy
This skill should be used when the user wants to "deploy an agent", "deploy my ADK agent", "set up CI/CD", "configure secrets", "troubleshoot a deployment", or needs guidance on Agent Runtime, Cloud Run, or GKE deployment targets, or binding an agent to an Agent Gateway. Covers deployment workflows,
By google · 191,565 installs
npx skills add google/agents-cli --skill google-agents-cli-deploy
Source repository · Upstream listing
Deployment Guide
Requires: agents cli ( uv tool install google agents cli ) — [install uv](https://docs.astral.sh/uv/getting started/installation/index.md) first if needed.
Prefer using the agents cli commands throughout this guide — they wrap Terraform, Docker, and deployment into a tested pipeline. If your project isn't scaffolded yet, see /google agents cli scaffold to add deployment support first.
Reference Files
For deeper details, consult these reference files in references/ :
cloud run.md — Scaling defaults, Dockerfile, session types, networking
agent runtime.md — container based deploy, unified FastAPI app, the /api passthrough, Terraform resource, deployment metadata, CI/CD differences
gke.md — GKE Autopilot cluster, Kubernetes manifests, Workload Identity, session types, networking
terraform patterns.md — Custom infrastructure, IAM, state management, importing resources
batch inference.md — BigQuery Remote Function trigger; for Pub/Sub / Eventarc on ADK see /google agents cli adk code
cicd pipeline.md — Full CI/CD pipeline setup, infra cicd flags, runner comparison, WIF auth, pipeline stages
testing deployed agents.md — Testing instructions per deployment target, curl examples, load tests
Observability: See the /google agents cli observability skill for Cloud Trace, prompt response logging, BigQuery Analytics, and third party integrations.
Deployment Target Decision Matrix
Choose the right deployment target based on your requirements:
Criteria Agent Runtime Cloud Run GKE
Scaling Managed auto scaling (configurable min/max, concurrency) Fully configurable (min/max instances, concurrency, CPU allocation) Full Kubernetes scaling (HPA, VPA, node auto provisioning)
Networking VPC SC and PSC I supported (private VPC connectivity via network attachments) Full VPC support, direct VPC egress, IAP, ingress rules Full Kubernetes networking
Session state Managed Agent Engine sessions (ADK wires VertexAiSessionService automatically) In memory (dev), Cloud SQL, or Agent Platform Sessions backend In memory (dev), Cloud SQL, or Agent Platform Sessions backend
Batch/event processing Trigger endpoints reachable via the Agent Engine /api passthrough Native trigger endpoints (Pub/Sub, Eventarc); ADK: see /google agents cli adk code Custom (Kubernetes Jobs, Pub/Sub)
Cost model vCPU hours + memory hours (not billed when idle) Per instance second + min instance costs Node pool costs (always on or auto provisioned)
Setup complexity Lower (managed, purpose built for agents) Medium (Dockerfile, Terraform, networking) Higher (Kubernetes expertise required)
Best for Managed infrastructure, minimal ops Custom infra, full networking control Full Kubernetes control
Ask the user which deployment target fits their needs. Each is a valid production choice with different trade offs.
All three targets are container based, so any language works.
Product name mapping: "Agent Engine" / "Vertex AI Agent Engine" is now Agent Runtime . Use deployment target agent runtime .
Ambient / scheduled / event driven agents (ADK projects): ADK's trigger sources registers /apps/{app}/trigger/ endpoints on the same FastAPI app for all targets. On Cloud Run / GKE these are public HTTP routes you point a Pub/Sub push subscription or Eventarc trigger at; on Agent Runtime the same routes are reachable through the Agent Engine /api passthrough (e.g. .../reasoningEngines/v1/{resource}/api/apps/{app}/trigger/pubsub ). Cloud Run remains the simplest target for unauthenticated trigger sources. See /google agents cli adk code ( references/adk python.md , section "12. Event Driven / Ambient Agents") for the trigger sources pattern.
OAuth / user consent agents: Use Agent Runtime with Gemini Enterprise for agents that need OAuth 2.0 user consent (e.g., accessing Google Drive, Calendar, or other user scoped APIs). Cloud Run does not currently support managed OAuth flows. For a worked ADK example, look up OAuth user consent in the topic index in /google agents cli adk code → references/samples.md .
Deploying to Dev
Deploy Workflow
Task tracking: Deployment involves multiple sequential steps (infra setup, CI/CD configuration, deploy, verification). Use a task list to track progress through these steps — skipping one often causes failures in later steps that are hard to trace back.
1. If prototype (no deployment target), first enhance: agents cli scaffold enhance . deployment target <target
2. Notify the human : paste the eval scores and test results, then ask "Ready to deploy to dev?"
3. Wait for explicit approval
4. Once approved: agents cli deploy
Agent Runtime timeout recovery: Agent Runtime deploys can take 5 10 minutes and may exceed command timeouts. If the deploy command is cancelled or times out, the deployment continues server side. Run agents cli deploy status to check progress — poll every 60 seconds until it reports completion or failure.
IMPORTANT : Never run agents cli deploy without explicit human approval.
Do NOT run agents cli infra single project before deploying. It is not a prerequisite — agents cli deploy works on its own. Run it separately if the user needs observability features (prompt response logging, BigQuery analytics) — see /google agents cli observability .
Single Project Infrastructure Setup (Optional — Advanced)
agents cli infra single project runs terraform apply in deployment/terraform/single project/ . Use this to provision single project GCP infrastructure without CI/CD (service accounts, IAM bindings, telemetry resources, Artifact Registry). Also useful to test things in a single project before going to production. It is NOT required for deploying.
Note: agents cli deploy doesn't automatically use the Terraform created app sa . Pass the service account explicitly: agents cli deploy service account SA EMAIL .
Deploy Flag Reference
Flag Description Targets
project GCP project ID All
region GCP region All
service account Service account email for the deployed agent All
service name Override the deployed service name (Cloud Run service or Agent Runtime display name); defaults to the project name. If you override it, consider updating your Terraform and CI (if present) — they name resources from the project name. Not supported for GKE, whose names are fully owned by Terraform. Agent Runtime, Cloud Run
secrets Comma separated ENV=SECRET or ENV=SECRET:VERSION pairs Agent Runtime, Cloud Run
update env vars Comma separated KEY=VALUE environment variables Agent Runtime, Cloud Run
agent identity Enable [agent identity](https://docs.cloud.google.com/gemini enterprise agent platform/scale/runtime/agent identity) (Preview) Agent Runtime
network attachment Network attachment resource name for [PSC interface](https://docs.cloud.google.com/gemini enterprise agent platform/scale/runtime/private service connect interface) (enables private VPC connectivity) Agent Runtime
dns peering domain DNS peering domain suffix, e.g. my internal.corp. (requires network attachment ) Agent Runtime
dns peering project Project ID hosting the Cloud DNS managed zone for DNS peering (requires network attachment ) Agent Runtime
dns peering network VPC network name in the target project for DNS peering (requires network attachment ) Agent Runtime
agent gateway egress Bind the agent to an [Agent Gateway](https://docs.cloud.google.com/gemini enterprise agent platform/govern/gateways/agent gateway overview) governing outbound traffic. Full resource name of a gateway with governedAccessPath=AGENT TO ANYWHERE . Empty value unbinds; omit to leave unchanged. See [Agent Gateway]( agent gateway) Agent Runtime
agent gateway ingress Bind the agent to an Agent Gateway governing inbound traffic. Full resource name of a gateway with governedAccessPath=CLIENT TO AGENT . Empty value unbinds; omit to leave unchanged Agent Runtime
memory Memory limit (default: 4Gi ) Agent Runtime, Cloud Run
cpu CPU limit (default: 1 ) Agent Runtime, Cloud Run
min instances Minimum number of instances (default: 0 , i.e. scale to zero; the generated Terraform uses 1 ) Agent Runtime, Cloud Run
max instances Maximum number of instances (default: 10 ) Agent Runtime, Cloud Run
concurrency Concurrent requests per container (default: 8 ; see [Sizing a deployment]( sizing a deployment)) Agent Runtime, Cloud Run
port Container port Cloud Run, Agent Runtime
build args Comma separated KEY=VALUE Docker build args Agent Runtime
labels Comma separated KEY=VALUE resource labels. Additive: adds/updates the labels you name; labels you don't name are preserved. Agent Runtime, Cloud Run
iap Enable Identity Aware Proxy Cloud Run
image Container image URI (skips source build; not supported for Agent Runtime) Cloud Run, GKE
no wait Start deployment and return immediately Agent Runtime, Cloud Run
status Check the status of a pending no wait deployment Agent Runtime, Cloud Run
list List existing deployments and exit All
dry run / n Print what would be executed without running it All
no confirm project Skip project confirmation prompt All
Run agents cli deploy help for the full flag reference.
Advanced Cloud Run Deploys: If you need features not exposed via agents cli flags, use dry run (or n ) to print the full gcloud command, copy it, and add additional arguments as needed.
Project Confirmation: If the project is resolved automatically (not passed via project ), the command will prompt for confirmation in interactive mode. Since agents typically run in non interactive mode, you MUST pass no confirm project to proceed if you are relying on automatic project resolution.
Sizing a deployment
Defaults (same on Agent Runtime and Cloud Run): cpu 1 , memory 4Gi , concurrency 8 , min instances 0 , max instances 10 . The generated service.tf matches, except it pins min instances = 1 so production deployments don't experience cold starts.
agents cli deploy scales to zero by default so idle dev and demo agents don't hold capacity. Pass min instances 1 (or deploy via Terraform) when you need a warm instance.
The params are coupled — scale them together:
One async process — scale out, not up. The container runs a single uvicorn process that serves many requests concurrently on the event loop, so throughput comes from concurrency and horizontal scale ( max instances ), not extra worker processes. Raise cpu only if profiling shows the event loop or synchronous tool calls are CPU bound.
Memory bounds concurrency. Each concurrent request keeps its full working set (context window, history, RAG chunks, response buffer) in memory while it waits on the model, so peak ≈ base + concurrency × per request memory . Memory — not CPU — is the first limit, so raising concurrency without memory is the main OOM cause.
Concurrency default is conservative. An async worker can serve many concurrent requests while it waits on the model, but per request memory is agent specific, so 8 protects a memory heavy (RAG/multimodal) agent. Light agents can raise it to 16–32+ after load testing. See [Underutilized asynchronous workers](https://docs.cloud.google.com/gemini enterprise agent platform/scale/runtime/optimize and scale underutilized workers).
Tune with the scaffolded load test ( tests/load test/ , run locally or in the CI/CD staging pipeline): drive load, watch max latency and memory/OOM restarts, then adjust — high max latency → raise concurrency (+ wo