Imagine controlling your web browser just by looking at screenshots. That is what this skill does. It uses Midscene Bridge mode to connect to your desktop Chrome browser. The AI agent sees everything on the screen and can click, type, scroll, or fill forms. It works with any website, no matter how it was built. Your login sessions and cookies stay safe. The mouse and keyboard are not taken over. Everything happens through Chrome DevTools Protocol.
This is useful for many tasks. You can browse pages, extract data, test new features, or automate repetitive steps. The skill is powered by Midscene.js and needs a visual AI model like Gemini or Qwen. Just set up the API key and model name. Then the agent can start working with your real browser.
Global
mkdir -p ~/.claude/skills/chrome-bridge-automationProject
mkdir -p .claude/skills/chrome-bridge-automationSource Repository
Agent Browservercel-labs/agent-browser
Fast browser automation tool for AI agents to control any website
Develop Userscriptsxixu-me/skills
Build and debug browser userscripts for Tampermonkey and ScriptCat
Just Scrapescrapegraphai/just-scrape
Search, scrape, crawl, and monitor web pages with an easy command line tool
Browser Actbrowser-act/skills
AI agents use browser-act to automate web tasks like fetch forms screenshots and complex workflows
Firecrawlfirecrawl/cli
Search, scrape, and crawl websites with one simple command line tool
Playwright Climicrosoft/playwright-cli
Automate browser tasks and test web pages with simple terminal commands
Browser Usebrowser-use/browser-use
Automate web testing, form filling, and screenshots with simple CLI commands
Browser Act Skill Forgebrowser-act/skills
Explore any website once and create reusable skill packages forever