An autonomous E2E testing agent that reads user stories, drives real browsers, and files evidence-backed bug reports β no test scripts required.
Traditional E2E testing breaks the moment a data-testid changes. Sentinel-QA doesn't rely on brittle selectors β it perceives the DOM the way a real user (or QA engineer) would: by role, by visible text, and by visual grounding when nothing else resolves.
Give it a user story or acceptance criteria. Sentinel plans the test, drives a real browser via Playwright, verifies the outcome, and β when something breaks β writes and files a reproducible bug report with video, trace, and console evidence attached automatically.
No hand-written test scripts. No selector maintenance. No "can't reproduce."
User Story β Test Plan β Browser Execution β Verification β Bug Report Filed
| Problem | How Sentinel solves it |
|---|---|
| Selectors break on every UI refactor | Semantic resolution: ARIA role/name β text match β vision-model grounding |
| Writing E2E tests is slow and someone has to maintain them | Tests are planned automatically from user stories / Gherkin acceptance criteria |
| Flaky tests erode trust in the suite | Built-in triage separates real bugs from flaky tests from selector drift |
| Bug reports without repro steps waste engineering time | Every filed bug ships with video, trace, and console/network evidence |
| Vendor lock-in to one LLM or one issue tracker | Pluggable LLM provider layer; pluggable Jira / GitHub Issues / Linear filers |
βββββββββββββββ ββββββββββββ βββββββββββββββ ββββββββββββββββ βββββββββββββ
β Ingestion β β β Planning β β β Perception β β β Execution β β βVerificationβ
β (user story)β β(step graph)β β(DOM resolve) β β (Playwright) β β (assertions)β
βββββββββββββββ ββββββββββββ βββββββββββββββ ββββββββββββββββ βββββββ¬ββββββ
β
βββββββββββββ ββββββββΌβββββββ
β Reporting β β β Triage β
β(file a bug)β β(classify fail)β
βββββββββββββ βββββββββββββββ
- Ingestion β parses a user story or Gherkin acceptance criteria into structured test intent
- Planning β an LLM generates an ordered, validated step graph (actions + assertions)
- Perception β resolves each step's target element semantically, not by fragile CSS selectors
- Execution β drives a real browser via Playwright with self-healing retries
- Verification β checks assertions, visual regressions, and console/network errors
- Triage β classifies any failure as a genuine bug, a flaky test, or selector drift
- Reporting β writes a structured bug report with attached evidence and files it to your tracker
# Clone and install
git clone https://github.com/morka17/qa-agent.git
cd qa-agent
poetry install
playwright install --with-deps chromium
# Configure
cp .env.example .env
# set your LLM provider key, target app URL, and issue tracker credentials in .env
# Run against a user story
poetry run qa_agent run \
--story "As a user, I can add an item to my cart and see the updated total" \
--target-url https://your-staging-app.com
# Or run against a full backlog (CSV / Jira / Linear)
poetry run qa_agent run --stories ./backlog.csvA successful run produces a report in runs/<run_id>/:
runs/2026-09-01-a1b2c3/
βββ plan.json # generated step graph
βββ trace.zip # Playwright trace (open with `playwright show-trace`)
βββ video.webm
βββ screenshots/
βββ report.md # filed bug report (if a failure was found)
Full system design, sequence diagrams, and Architecture Decision Records live in docs/architecture.md and docs/adr/. Key design principles:
- Layered selector resolution β ARIA/text matching first; LLM reasoning and vision-model grounding are fallbacks, not the default (for cost and reliability)
- Versioned, auditable prompts β every LLM call traces back to a reviewed prompt spec in
docs/prompts/, not an inlined string - Evidence-first bug reports β no report is filed without attached trace/video/screenshot proof
- Replayable runs β every run can be replayed via
scripts/replay_trace.pyfor debugging - Guardrails before autonomy β destructive actions (delete, payment, irreversible submit) require an allowlist match or human approval gate
- Core perception + execution loop against fixture apps
- Story β test plan pipeline
- Failure triage & root-cause analysis
- Automated bug filing (Jira / GitHub Issues / Linear)
- Multi-app concurrent orchestration + operator dashboard
- Observability (OpenTelemetry, cost tracking) & guardrails hardening
- Public benchmark suite (selector accuracy, flake rate, cost per run)
See the full roadmap for phase-by-phase detail.
Sentinel-QA is open source and contributions are welcome β whether that's a bug fix, a new selector strategy, a new issue-tracker integration, or a benchmark result.
- Read CONTRIBUTING.md
- Check open issues labeled
good first issue - Run
make setup && make testbefore opening a PR
Python Β· FastAPI Β· Playwright Β· PostgreSQL + pgvector Β· SQLAlchemy (async) Β· Alembic Β· Redis/Temporal Β· Docker Β· Terraform Β· OpenTelemetry
MIT β free to use, modify, and self-host.
Built for teams who'd rather review a filed bug report than write another Playwright script by hand.