webbrain-one/webbrain

β–² 18 stars todayβ˜… 1,183β‘‚ 133

Open-source AI browser agent for Chrome and Firefox (monorepo) 🧠

About webbrain-one/webbrain

webbrain-one/webbrain is an open-source project on GitHub, mainly written in JavaScript. Open-source AI browser agent for Chrome and Firefox (monorepo) 🧠 It currently holds 1,183 stars and 133 forks with 0 open issues, and was last pushed on an unknown date (repository created unknown).

Project Overview

AI Homed tracks it on the Today's Trending board, currently at rank #51 with 18 new stars today.

GitHub Repository Details

Repository webbrain-one/webbrain Β· default branch - Β· size 0 KB Β· watchers 0 Β· source: GitHub REST API and repository README

README

https://github.com/webbrain-one/webbrain/blob/HEAD/WebBrain logo

WebBrain

Open-source AI browser agent for chatting with pages, automating tasks, and running multi-step workflows with your choice of LLM.

https://github.com/webbrain-one/webbrain/blob/HEAD/Install WebBrain from the Chrome Web Store https://github.com/webbrain-one/webbrain/blob/HEAD/Install WebBrain from Firefox Browser Add-ons https://github.com/webbrain-one/webbrain/blob/HEAD/Install WebBrain from Microsoft Edge Add-ons

English Β· δΈ­ζ–‡ Β· FranΓ§ais Β· Docs Β· Website Β· Discord Β· GPL-3.0-or-later

WebBrain reading a page, filling in a form, and fetching a file

WebBrain is a web browser extension that puts an AI agent in a side panel next to your tabs. Ask it about the page you're on, or hand it a task and let it click, type, and navigate its way through. It runs on the model you choose β€” a local llama.cpp or Ollama server, a frontier cloud API, or the managed default that needs no setup at all.

Install

Install from the Chrome Web Store, Firefox Add-ons, or Edge Add-ons.

Or load it from source
git clone https://github.com/webbrain-one/webbrain.git

Chrome β€” open chrome://extensions/, enable Developer mode (top right), click Load unpacked, and select the webbrain/src/chrome folder. For an isolated copy that won't pick up stray working-tree files, run npm run build:chrome first and load webbrain/build/chrome instead.

Firefox β€” open about:debugging#/runtime/this-firefox, click Load Temporary Add-on, and select src/firefox/manifest.json (or build/firefox/manifest.json after npm run build:firefox). Temporary add-ons are removed when Firefox restarts; permanent installation requires signing via addons.mozilla.org.

Use it

Click the WebBrain icon to open the side panel, then type something like:

Three modes control what the agent is allowed to do:

| Mode | What it can do | | ------- | ---------------------------------------------------------------------- | | Ask | Read-only. Reads the page, answers questions, fetches URLs. | | Act | Clicks, types, navigates, uploads, downloads, fills forms. | | Dev | Adds page source, styles, console, network, and reversible page edits. |

Pick a model

WebBrain Compass 1.0 is the default and needs no API key or local setup.

Local models need no API key either. Point WebBrain at any OpenAI-compatible server:

llama-server -m your-model.gguf --port 8080          # llama.cpp
ollama serve                                          # Ollama  β†’ :11434/v1
vllm serve your-model --port 8000                     # vLLM    β†’ :8000/v1
python -m sglang.launch_server --model-path your-model --port 30000

LM Studio (:1234/v1), Osaurus (:1337/v1), Jan (:1337/v1), LocalAI (:8080/v1), and GPT4All (:4891/v1) work the same way. A generic Local OpenAI-compatible Proxy card also supports authenticated loopback gateways such as CLIProxyAPI; see the secure subscription proxy setup. Unsloth Studio (Local) uses the same OpenAI-compatible path with the Studio port and sk-unsloth- API key configured by the user; see the Unsloth Studio setup. Osaurus (Local) connects to the Mac's Osaurus server with model discovery, streaming, and tool calls; see the Osaurus setup. Load a model with at least a 16k-token context window β€” 8k works only with the Compact tier, and 4k is too small for the system prompt plus tool schemas. WebBrain auto-detects the real window for llama.cpp, Ollama, and LM Studio, and auto-compacts the conversation as it fills up. For Ollama, llama.cpp, LM Studio, and LocalAI, it also reads native server metadata before adding screenshots; Settings provides Auto, Force on, and Off overrides. When the optional Model field is blank, the loaded-model capability is rechecked on every user turn so a server-side hot swap takes effect. There is also a preview ollama launch webbrain --model handoff. Details: providers and models.

Cloud APIs β€” OpenAI, Anthropic Claude, Google Gemini, Azure OpenAI, AWS Bedrock, Mistral, DeepSeek, xAI Grok, MiniMax, Kimi, Qwen, z.ai GLM, Groq, Together, Cloudflare, Nvidia NIM, Hugging Face, Fireworks, OpenRouter, and more. Settings ships 112 built-in provider cards on Chromium (111 on Firefox), including an endpoint-free local WebGPU option with the tested LFM2.5 2.6B preset and an experimental custom Hugging Face ONNX repository option β€” see the full catalog.

For a local or bring-your-own provider, the per-provider Share queries for research switch remains off by default. When enabled, it now shares a bounded, content-free diagnostic timeline for that provider's model attempts (including failed runs) alongside the existing scrubbed prompt/response share. The timeline includes tool names, outcomes, error codes, and timings, but not tool arguments, page content, or screenshots. If a run fails before a normal generation share, its bounded model-facing request and final blocker accompany the diagnostic record. No second sharing switch is required; turning the existing switch off also purges queued diagnostics.

Features

elements, via the accessibility tree rather than brittle selectors forms, with per-site permission prompts before consequential actions approval, and pin the approved plan to the scratchpad before any tool runs (default 130), with a Continue button when it hits the limit workflow you can re-run, export, and share page and act when a condition is met emergency overflow recovery user memory for stated preferences keep your question in view as answers grow, copy buttons, a page-inspection banner, and a stop button that works mid-run decisions, 0.3 for Ask, 0 for vision screenshot descriptions

Agent tools

WebBrain separates tier from mode. Tier (compact, mid, full) is a per-provider setting controlling how many tools a model sees β€” Compact suits small local models, Full unlocks hover, drag-drop, frames, and shadow DOM. Mode (ask, act, dev) controls what the user is allowing.

The full tool-by-tier matrix, WebMCP notes, and Dev-mode diagnostics are in agent tools.

Slash commands

Type /help in the panel for full signatures and flags. The most useful ones:

| Command | What it does | | -------------------------------------------------------------------------------------- | --------------------------------------------------------------------------- | | /ask Β· /act Β· /dev Β· /plan | Switch mode before sending | | /schedule [prompt] | Create a scheduled task | | /watch [--keep] [--secs <30-120>] [--long \| --short] [/beep] | Poll the current page and act when a condition is met | | /workflow Β· /workflow --save | Manage saved workflows or compile the last successful run | | /teach --start Β· /teach --end | Learn a reusable workflow from your demonstrated actions | | /memory --add | Save a user preference | | /screenshot [--full-page] | Capture the tab, or the full scrollable page | | /record [--transcribe] | Record the current tab, optionally saving a transcript | | /export [--traces \| --config] | Download the conversation, tool chain, or a Settings snapshot | | /compact Β· /reset Β· /verbose | Compact context, clear the conversation, toggle tool detail | | /allow-api | Per-conversation override letting fetch_url mutate when the UI is failing |

/watch runs its first check immediately, then polls every 60 seconds (--secs accepts 30–120). Relative conditions like "when a new commit appears" establish a baseline on the first check; --keep keeps the watch running and suppresses repeated alerts for the same stable event key.

Full reference, including /dangerously-skip-permissions and the run-capture suffixes: slash commands.

Keyboard Shortcuts

Chrome side panel shortcuts work when the WebBrain side panel has focus.

| Shortcut | What it does | | ------------------------------- | ---------------------------------------------------------------------------- | | Ctrl+/ or Cmd+/ | Focus the input | | Ctrl+Shift+A or Cmd+Shift+A | Switch to Ask mode | | Ctrl+Shift+X or Cmd+Shift+X | Switch to Act mode | | Ctrl+Shift+D or Cmd+Shift+D | Switch to Dev mode | | Escape | Stop the active run, unless it is only dismissing slash-command autocomplete | | Escape twice | Stop an active recording from WebBrain or browser pages |

Documentation

| | | | ------------------------------------------------------------------------------------------------------------------------ | -------------------------------------------------------- | | Architecture | System overview, turn flow, subsystems | | Agent tools | Tiers, modes, and the full tool matrix | | Slash commands | Every command and flag | | Providers and models | All provider cards, local setup, tiers | | Skills | Bundled skills, importing, skill tools | | Security model | Permissions, credentials, trust boundaries | | Prompt-injection defense | Defense layers and known gaps | | Privacy and data flow | What leaves the browser, and what doesn't | | Accessibility tree and refs | How pages are read and targeted | | Site adapters | Per-site guidance and versioned workflow contracts | | Export and workflow formats | webbrain-config/1, webbrain-workflow/1 | | Adding a tool Β· Localization Β· Test scenarios | Contributor guides | | Community | Discord server guide: channels, roles, rules, escalation |

Also available in δΈ­ζ–‡ and FranΓ§ais.

Community

Chat about everything WebBrain β€” help, local and cloud model setups, site adapters, show-and-tell, and contributor coordination β€” on the WebBrain Discord. See community for how the server is organized, and discord-setup for the channel, role, and welcome-screen configuration. Bug reports and feature requests belong in GitHub issues, not Discord.

Repository layout

src/chrome/     Manifest V3 build β€” service worker, chrome.scripting, sidePanel
src/firefox/    Manifest V2 build β€” background page, executeScript, sidebar_action
docs/           Design and reference docs (en, zh-CN, fr)
mcp-server/     MCP server β€” delegate browser tasks from Claude Code, Codex, Cursor
lmstudio-plugin/  Web tools + browser delegation as a standalone LM Studio plugin
web/            Landing site and docs site
test/           Node test suite, LLM scenario benchmarks, security corpora

Nearly all agent code is shared between the two builds. See architecture for where they diverge.

Known issues

Firefox is meaningfully weaker than Chrome. Firefox has no equivalent to the Chrome DevTools Protocol via chrome.debugger, so the Firefox build has no shadow-DOM piercing, no real trusted mouse events (some React/Vue handlers won't fire), no closed-shadow-root traversal, no resolveSelector retry budget, no SPA-navigation-aware retry, and no CDP screenshots. It uses tabs.captureTab for viewport screenshots, including inactive run tabs, but still cannot provide Chrome's pixel-perfect or full-page CDP capture. Site adapters, vision detection, loop detection, the auto-screenshot loop, and the Compact prompt/tool set _are_ mirrored to Firefox. Some single-page apps may also fail to trigger content-script re-injection after client-side navigation.

Contributing

See CONTRIBUTING.md. To add a tool, follow the checklist in adding a tool. To add a provider, subclass BaseLLMProvider, implement chat() (and optionally chatStream()), and register it in providers/manager.js β€” mirroring both changes to src/chrome/ and src/firefox/. All providers normalize to { content, toolCalls, usage }; details in providers and models.

Recent changes are in CHANGELOG.md.

MCP server

Let a coding agent use your browser. Claude Code, Codex, Cursor and OpenClaw can delegate a task to WebBrain running in the session you are already signed into β€” cookies present, SSO already passed. A headless framework starts logged out and stalls at the first login wall; this does not.

claude mcp add --transport stdio webbrain -- npx -y @webbrain/mcp-server

Claude Code launches the server automatically when it starts an MCP session. To launch it yourself instead, run the following command and leave that terminal open (press Ctrl+C to stop it):

npx -y @webbrain/mcp-server

Once the server is running, open WebBrain β†’ Settings β†’ General β†’ Advanced β†’ MCP, set the URL to ws://127.0.0.1:17374/extension, and enable it. Chromium only β€” the control and bridge runtime use the extension's off-screen document, which the Firefox build does not have.

If Settings reports Connection error: WebSocket error, nothing is normally listening at the configured URL. Start the MCP server, confirm that the URL uses port 17374, and leave its process running. See the mcp-server troubleshooting guide for a listener check and the other bridge ports.

webbrain_run(task: "open the Stripe dashboard and list last week's failed
             payments with amounts and customer emails", mode: "ask")

Use webbrain_extract with a JSON Schema when the caller needs predictable structured output instead of a prose summary. The server exposes six task-level tools: run, structured extraction, status, clarification response, abort, and connection diagnostics.

mode='ask' is read-only. mode='act' can click and type, gated by the same in-browser approval prompts a human gets. The server exposes task delegation rather than the ~50 low-level browser primitives: WebBrain's permission gate lives in the agent loop, so per-primitive access over a socket would sit below the gate and bypass it. Details in mcp-server/.

The complete client setup, tool arguments, run lifecycle, structured-output examples, safety boundaries, and troubleshooting guide live at web/docs/mcp/.

The extension holds one bridge socket at a time β€” WebBrain Cloud (17373),
the MCP server (17374), or the LM Studio plugin (17375). Switch by changing
the URL under Settings β†’ General β†’ Advanced β†’ MCP.

LM Studio plugin

A standalone LM Studio plugin at webbrain/web-tools:

lms clone webbrain/web-tools

fetch_url and research_url are pure Node HTTP β€” no browser needed, but also no cookies, no session and no JavaScript. With the extension installed on a Chromium browser, browser_task adds delegation to your real signed-in browser, reaching the authenticated and client-rendered pages plain HTTP cannot. It degrades with an actionable message when no extension is attached, and the HTTP tools keep working on Firefox.

Source: lmstudio-plugin/.

Contributors

Citation

@software{webbrain2026,
  author = {Sokullu, Emre},
  title = {WebBrain: Open-source AI browser agent for chatting with pages},
  year = {2026},
  publisher = {GitHub},
  url = {https://github.com/webbrain-one/webbrain}
}

License

WebBrain is licensed under GPL-3.0-or-later because the distributed browser extension bundles and integrates the GPL-licensed Xapian/libzim WebAssembly runtime.

Built with ❀️ by Emre Sokullu and open-source contributors.

GitHub Stars & Activity

1,183Stars
133Forks
0Open issues
JavaScriptLanguage

GitHub Popularity

GitHub stars1,183
Forks133
Open issues0
Primary languageJavaScript
License-
Stars gained today18
Created-
Last pushed-

Trending History

Daily boardrank #51 Β· β–² 18 stars

Related AI Projects

1

affaan-m / ECC

JavaScriptβ˜… 271,104β‘‚ 40,507β–² 572 stars
β†’
2

DietrichGebert / ponytail

JavaScriptβ˜… 151,548β‘‚ 8,127β–² 1,429 stars
β†’
3

addyosmani / agent-skills

JavaScriptβ˜… 100,478β‘‚ 10,556β–² 188 stars
β†’
4

pbakaus / impeccable

JavaScriptβ˜… 74,193β‘‚ 4,471β–² 717 stars
β†’
5

coreyhaines31 / marketingskills

JavaScriptβ˜… 52,344β‘‚ 7,875β–² 139 stars
β†’
6

vercel-labs / agent-skills

JavaScriptβ˜… 31,847β‘‚ 2,787β–² 55 stars
β†’
7

mnfst / awesome-free-llm-apis

JavaScriptβ˜… 8,996β‘‚ 922β–² 146 stars
β†’
8

BuilderIO / skills

JavaScriptβ˜… 4,511β‘‚ 227β–² 14 stars
β†’

More AI Rankings