Skip to content

Latest commit

Β 

History

56 Commits

Folders and files

NameName
Last commit message
Last commit date
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 
Β 

Repository files navigation

πŸ›‘οΈ Sentinel-QA

An autonomous E2E testing agent that reads user stories, drives real browsers, and files evidence-backed bug reports β€” no test scripts required.

CI License Python Playwright PRs Welcome Discord


What is Sentinel-QA?

Traditional E2E testing breaks the moment a data-testid changes. Sentinel-QA doesn't rely on brittle selectors β€” it perceives the DOM the way a real user (or QA engineer) would: by role, by visible text, and by visual grounding when nothing else resolves.

Give it a user story or acceptance criteria. Sentinel plans the test, drives a real browser via Playwright, verifies the outcome, and β€” when something breaks β€” writes and files a reproducible bug report with video, trace, and console evidence attached automatically.

No hand-written test scripts. No selector maintenance. No "can't reproduce."

User Story  β†’  Test Plan  β†’  Browser Execution  β†’  Verification  β†’  Bug Report Filed

Why Sentinel-QA

Problem How Sentinel solves it
Selectors break on every UI refactor Semantic resolution: ARIA role/name β†’ text match β†’ vision-model grounding
Writing E2E tests is slow and someone has to maintain them Tests are planned automatically from user stories / Gherkin acceptance criteria
Flaky tests erode trust in the suite Built-in triage separates real bugs from flaky tests from selector drift
Bug reports without repro steps waste engineering time Every filed bug ships with video, trace, and console/network evidence
Vendor lock-in to one LLM or one issue tracker Pluggable LLM provider layer; pluggable Jira / GitHub Issues / Linear filers

How It Works

β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”
β”‚  Ingestion  β”‚ β†’  β”‚ Planning β”‚ β†’  β”‚ Perception   β”‚ β†’  β”‚  Execution   β”‚ β†’  β”‚Verificationβ”‚
β”‚ (user story)β”‚    β”‚(step graph)β”‚  β”‚(DOM resolve) β”‚    β”‚ (Playwright) β”‚    β”‚ (assertions)β”‚
β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”¬β”€β”€β”€β”€β”€β”˜
                                                                                  β”‚
                                                          β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”    β”Œβ”€β”€β”€β”€β”€β”€β–Όβ”€β”€β”€β”€β”€β”€β”
                                                          β”‚  Reporting β”‚ ←  β”‚   Triage    β”‚
                                                          β”‚(file a bug)β”‚    β”‚(classify fail)β”‚
                                                          β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜    β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜
  1. Ingestion β€” parses a user story or Gherkin acceptance criteria into structured test intent
  2. Planning β€” an LLM generates an ordered, validated step graph (actions + assertions)
  3. Perception β€” resolves each step's target element semantically, not by fragile CSS selectors
  4. Execution β€” drives a real browser via Playwright with self-healing retries
  5. Verification β€” checks assertions, visual regressions, and console/network errors
  6. Triage β€” classifies any failure as a genuine bug, a flaky test, or selector drift
  7. Reporting β€” writes a structured bug report with attached evidence and files it to your tracker

Quickstart

# Clone and install
git clone https://github.com/morka17/qa-agent.git
cd qa-agent
poetry install
playwright install --with-deps chromium

# Configure
cp .env.example .env
# set your LLM provider key, target app URL, and issue tracker credentials in .env

# Run against a user story
poetry run qa_agent run \
  --story "As a user, I can add an item to my cart and see the updated total" \
  --target-url https://your-staging-app.com

# Or run against a full backlog (CSV / Jira / Linear)
poetry run qa_agent run --stories ./backlog.csv

A successful run produces a report in runs/<run_id>/:

runs/2026-09-01-a1b2c3/
β”œβ”€β”€ plan.json          # generated step graph
β”œβ”€β”€ trace.zip          # Playwright trace (open with `playwright show-trace`)
β”œβ”€β”€ video.webm
β”œβ”€β”€ screenshots/
└── report.md          # filed bug report (if a failure was found)

Architecture

Full system design, sequence diagrams, and Architecture Decision Records live in docs/architecture.md and docs/adr/. Key design principles:

  • Layered selector resolution β€” ARIA/text matching first; LLM reasoning and vision-model grounding are fallbacks, not the default (for cost and reliability)
  • Versioned, auditable prompts β€” every LLM call traces back to a reviewed prompt spec in docs/prompts/, not an inlined string
  • Evidence-first bug reports β€” no report is filed without attached trace/video/screenshot proof
  • Replayable runs β€” every run can be replayed via scripts/replay_trace.py for debugging
  • Guardrails before autonomy β€” destructive actions (delete, payment, irreversible submit) require an allowlist match or human approval gate

Roadmap

  • Core perception + execution loop against fixture apps
  • Story β†’ test plan pipeline
  • Failure triage & root-cause analysis
  • Automated bug filing (Jira / GitHub Issues / Linear)
  • Multi-app concurrent orchestration + operator dashboard
  • Observability (OpenTelemetry, cost tracking) & guardrails hardening
  • Public benchmark suite (selector accuracy, flake rate, cost per run)

See the full roadmap for phase-by-phase detail.


Contributing

Sentinel-QA is open source and contributions are welcome β€” whether that's a bug fix, a new selector strategy, a new issue-tracker integration, or a benchmark result.

  1. Read CONTRIBUTING.md
  2. Check open issues labeled good first issue
  3. Run make setup && make test before opening a PR

Tech Stack

Python Β· FastAPI Β· Playwright Β· PostgreSQL + pgvector Β· SQLAlchemy (async) Β· Alembic Β· Redis/Temporal Β· Docker Β· Terraform Β· OpenTelemetry


License

MIT β€” free to use, modify, and self-host.


Built for teams who'd rather review a filed bug report than write another Playwright script by hand.

About

πŸ›‘οΈ Sentinel-QA - An autonomous E2E testing agent that reads user stories, drives real browsers via Playwright, self-resolves DOM elements with LLM+heuristic reasoning, and files evidence-backed bug reports. No brittle selectors. No manual test scripts. Just ship, and let Sentinel watch the app.

Topics

Resources

Contributing

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages