gemini-computer-use
Build and run Gemini 2.5 Computer Use browser-control agents with Playwright. Use when a user wants to automate web browser tasks via the Gemini Computer Use model, needs an agent loop (screenshot → function_call → action → function_response), or asks to integrate safety confirmation for risky UI ac
By am-will · 1,159 installs
npx skills add am-will/codex-skills --skill gemini-computer-use
Source repository · Upstream listing
Gemini Computer Use
Quick start
1. Source the env file and set your API key:
2. Create a virtual environment and install dependencies:
3. Run the agent script with a prompt:
Browser selection
Default: Playwright's bundled Chromium (no env vars required).
Choose a channel (Chrome/Edge) with COMPUTER USE BROWSER CHANNEL .
Use a custom Chromium based executable (e.g., Brave) with COMPUTER USE BROWSER EXECUTABLE .
If both are set, COMPUTER USE BROWSER EXECUTABLE takes precedence.
Core workflow (agent loop)
1. Capture a screenshot and send the user goal + screenshot to the model.
2. Parse function call actions in the response.
3. Execute each action in Playwright.
4. If a safety decision is require confirmation , prompt the user before executing.
5. Send function response objects containing the latest URL + screenshot.
6. Repeat until the model returns only text (no actions) or you hit the turn limit.
Operational guidance
Run in a sandboxed browser profile or container.
Use exclude to block risky actions you do not want the model to take.
Keep the viewport at 1440x900 unless you have a reason to change it.
Resources
Script: scripts/computer use agent.py
Reference notes: references/google computer use.md
Env template: env.example