Let AI work for you in the browser. Just tell it what to do — WebSpider handles the rest.
WebSpider is a Chrome side-panel agent that drives your real browser: it reads the page, understands the controls, clicks, types, extracts, audits, verifies its own work and recovers when a model fails. No separate software, no cloud dashboard — one extension, your own API key.
English · فارسی
- Why WebSpider
- Features
- Install
- Quick start
- Connect a model
- Tools the agent can use
- Modes
- Speed and accuracy
- Human verification and CAPTCHA
- Privacy
- Keyboard
- Requirements
- Architecture
- Verification
- Troubleshooting
- Contributing
- License
- Creator and support
Most "AI browser" tools are a chat box bolted onto a page. They cannot see the controls, they guess CSS selectors, and when a click silently does nothing they happily report success. WebSpider was built around the opposite assumption: an agent that cannot prove what it did is not an agent.
So every control gets a stable elementId, every mutating action is followed by a real
page-state read, and the result tells you whether the page actually changed. When it did
not, WebSpider says so instead of counting it as a win.
| 🧰 45 built-in tools | Read, act, wait, audit and safety tools — see the catalogue or Settings → Tools in the panel. |
| ✅ Verifies its own actions | Every mutating action is followed by a cheap page-state read; the answer includes whether the page actually changed. |
| 🧠 Three modes | Agent (acts), Plan (plans first and cannot mutate the page), Research (read-only evidence). |
| 🔎 Deterministic audits | SEO, accessibility (real WCAG contrast maths), performance, links and content — each scored 0–100 with a severity-ranked issue list. |
| ⚡ Fast by default | Read-only tool calls run concurrently, snapshots are cached, and three speed profiles trade thoroughness for round trips. |
| 🛡️ Human-verification aware | Detects a CAPTCHA, pauses, shows you the screen, and resumes from the same step the moment you solve it. It never solves or bypasses one. |
| 🔁 Model failover | A timeout on one model automatically tries the next known-good one. Auth errors never trigger failover. |
| 🌊 Live streaming | Token-by-token answers, with an automatic non-streaming retry when a provider rejects stream:true. |
| ✋ Confirmation gate | Purchases, deletion, sending, submission and other sensitive actions pause for your approval first. |
| 🖱️ Stable element targeting | Location falls back through elementId → selector → role → text → index, folding case, punctuation, curly quotes and diacritics. |
| 🪟 Cross-frame aware | Works in iframes and shadow DOM, and re-targets itself when a click opens a popup, OAuth window or new tab. |
| 🗂️ Console & dialog capture | Reads the page's own console.* output and native alert/confirm/prompt calls via a MAIN-world bridge. |
| 💬 Full interface | Dark/light/system theme, four accents, three text sizes, Markdown answers with copy buttons, chat history, JSON export, toasts, real modals, keyboard shortcuts, RTL and reduced-motion support. |
- Open
chrome://extensionsand turn on Developer mode. - Click Load unpacked and select this folder.
- Pin WebSpider to the toolbar, then click it (or press
Alt+Shift+W) to open the side panel. - Open Settings → Connection, enter an OpenAI-compatible base URL and API key, then press Save & test.
There is nothing to build and no dependencies to install — the extension is plain HTML, CSS and JavaScript.
- Open any page you want to work with.
- In the panel, type what you want in plain language — for example "Research this page, find the main issues and write a short summary."
- Press Send. Watch the steps stream in as WebSpider inspects, acts and verifies.
- Approve any confirmation prompt, and complete any CAPTCHA yourself — the run resumes automatically the moment you are done.
WebSpider speaks the OpenAI Chat Completions protocol, so anything compatible works.
| Base URL | e.g. https://api.openai.com/v1 — WebSpider resolves /chat/completions and /models from it. |
| API key | Stored only in chrome.storage.local on your machine. Never logged, never shown in an error. |
| Works with | OpenAI, Azure OpenAI, OpenRouter, Groq, DeepSeek, Together, Mistral, Ollama, LM Studio, vLLM, and any other OpenAI-compatible gateway. |
Press Save & test and WebSpider lists the available models and health-checks them,
reporting latency. If /models is unavailable, pin a model id manually in the same tab.
Your API key is sent only to the endpoint you configured.
Read & inspect
| Tool | What it returns |
|---|---|
page_snapshot |
Structured page overview with stable elementIds for controls |
page_state |
Cheap state read plus a content hash, for "did anything change?" |
page_diagnostics |
Page errors, console errors, broken images, slow resources, suspicious links |
query_elements |
Structured list for any CSS selector (role, name, text, rect, visibility) |
get_table |
Headers, rows, or objects keyed by header |
get_form |
Fields with labels, values, required flags, options and validation state |
find_in_page |
Text matches with surrounding context, and the elements containing them |
get_selection |
The text the user has selected |
console_logs |
Console messages captured in the page's own JavaScript world |
dialog_log |
Native alert / confirm / prompt calls the page raised |
network_log |
Resource timings, transfer sizes, cache status |
styles_of |
Computed styles, effective background, WCAG contrast ratio |
audit |
seo · a11y · perf · links · content · all, scored and ranked |
extract |
Readable text, HTML, links or a table from an area |
mark_page / diff_page |
Prove exactly what changed |
Act on the page — navigate · open_tab · back · forward · click · click_at ·
type · clear_field · select · check · hover · focus · press_key ·
submit_form · scroll · scroll_to · highlight · draw · screenshot
Wait & verify — wait · wait_for_element · wait_for_text · wait_for_dom_stable ·
wait_for_network_idle
Safety & control — detect_human_check · human_verification · ask_confirmation ·
set_dialog_policy · sync_browser_target
Snapshots label every control with a stable elementId; the agent targets those first,
which is far more reliable than guessing CSS selectors.
| Mode | What it does |
|---|---|
| Agent | Full capability — reads the page and acts on it. |
| Plan | Investigates and returns a step-by-step plan. First-turn tool calls are discarded, so a plan can never change the page. |
| Research | Read-only evidence gathering. Acts are refused; the agent reads, extracts and audits only. |
| Setting | Default | Effect |
|---|---|---|
| Execution speed | Balanced | Fast skips post-action verification and caches longer; Thorough never caches and re-reads after every action |
| Verify the page after each action | On | Attaches an after block to every mutating action |
| Reuse a recent page snapshot | On | Saves a round trip when the agent inspects repeatedly |
| Pause for CAPTCHA, resume automatically | On | Never solves a challenge; pauses and resumes |
Read-only tool calls issued in the same turn run concurrently. Any mutating call is a barrier, so the order the model asked for is the order the page sees.
WebSpider detects human-verification challenges — reCAPTCHA, hCaptcha, Cloudflare Turnstile, Arkose, GeeTest, DataDome, PerimeterX, AWS WAF and MTCaptcha — names the vendor, pauses the run and shows you the screen. It then polls the page and resumes from the same step the instant you have solved it.
It will never solve, click, interpret or bypass a challenge. That is a deliberate design decision, not a missing feature: a CAPTCHA is a security control, and automating it enables mass account creation and abuse.
What happens on your screen stays on your machine.
- No telemetry, no analytics, no accounts, no auto-update pings.
- Page content is read locally. The only bytes that leave your computer are the ones the model needs to answer your request, sent only to the endpoint you configured.
- Your API key lives in
chrome.storage.localand is never logged or echoed. - Your conversations are stored locally as
chatSessionsand are yours to export or delete from the History drawer.
| Shortcut | Action |
|---|---|
Enter |
Send (configurable in Settings → Agent) |
Shift+Enter |
New line |
Ctrl/Cmd+K |
Focus the composer |
Alt+1 / 2 / 3 |
Agent / Plan / Research |
Esc |
Close the modal or drawer |
Alt+Shift+W |
Open the WebSpider side panel |
To run: Chrome 114 or newer (Manifest V3, side panel API). Nothing else — no build step, no dependencies, no background service.
To develop: any text editor. To run the test suite you need Node 22+ and jsdom.
| File | Role |
|---|---|
manifest.json |
MV3 manifest: side panel, service worker, content script, MAIN-world bridge |
main.js |
MAIN-world bridge (document_start): console + native dialog capture |
sidepanel.html |
Panel markup and the SVG icon sprite |
style.css |
Design tokens, both themes, logical-property layout |
sidepanel.js |
UI state machine, rendering, i18n, tool catalogue, persistence, export |
background.js |
Agent loop, SSE streaming, model failover, tool dispatch, verification |
content.js |
Page agent: locate, snapshot, readers, audits, actions, waits, overlays |
Conversations are stored in chrome.storage.local as chatSessions; the live agent state
is agentState. Storage keys are part of the data contract — renaming one loses history.
content.js runs in the isolated world and cannot see the page's JavaScript. main.js runs
in the page's own world and passes what it captures back through DOM slots
(#__webspider_console, #__webspider_dialogs) — the two worlds share the DOM, not the
JavaScript heap.
NODE_PATH=<node-workspace>/node_modules node .workbuddy-ai/tools/run-suite.mjs
Four stages, all run in-process:
| Stage | What it proves |
|---|---|
check-syntax.mjs |
Every file parses; icons, i18n keys and DOM selectors resolve; tool wiring is complete; the manifest is valid |
test-core.mjs |
Endpoint parsing, message-trim invariants, SSE assembly, streaming fallback, tool schemas, speed profiles, the human-check watchdog, and a whole run driven end-to-end against a stubbed model |
test-panel.mjs |
The real panel in jsdom: boot, quick actions, modes, i18n, theme, streaming, persistence, delete flows, keyboard, the tool catalogue, execution settings, export |
test-content.mjs |
The real content script in jsdom: snapshot ids, actions, waits, readers, audits, human-check detection, dispatch hardening |
661 assertions plus 12 static checks, all green. The static stage carries four
cross-file guards: worker dispatch ↔ content-script handler, declared tool ↔ act()
branch, declared tool ↔ panel catalogue, and the MAIN-world manifest registration.
Every guard in the suite was validated by injecting the bug it exists to catch — a guard that stays green under its own bug is decoration, not a test.
| Symptom | Fix |
|---|---|
| "API Key is missing" | Open Settings → Connection, paste a key and press Save & test. |
| "All verified models and endpoints failed" | Check the base URL, the key and your network. The panel shows the last error per model. |
| The agent reports the page "did NOT change" | The click landed on nothing. Ask it to take a snapshot or run page_diagnostics. |
| A run pauses and shows a challenge | That is the CAPTCHA gate. Solve it in the page — the run resumes by itself. |
| The side panel does not open | Pin the extension and click its icon, or press Alt+Shift+W. |
Issues and pull requests are welcome — see CONTRIBUTING.md for the workflow and house rules. To report a vulnerability privately, follow SECURITY.md.
MIT.
Built by Reza Kazemi Fathi.
If WebSpider saved you some clicking, a ⭐ on the repository genuinely helps.
Support the project: Daramet (IRR) · Donatr.ee (USD / crypto)
