Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
39 commits
Select commit Hold shift + click to select a range
5215af0
Add verified discovery recording and monochrome replay foundation
copyleftdev Sep 21, 2026
4e2f15f
Reserve shared demo budget before dispatch and retain uncertain charges
copyleftdev Sep 21, 2026
14034b0
Ingest bounded real TREC corpus samples with verified evidence locations
copyleftdev Sep 21, 2026
44a4923
Validate source-backed review findings and prepare private review tasks
copyleftdev Sep 21, 2026
ab7f2fe
Execute bounded reviewer tasks through Braess with validated evidence…
copyleftdev Sep 21, 2026
62f3800
Inventory real native attachments and retain bounded OCR page evidence
copyleftdev Sep 21, 2026
4774f3c
Bind OCR review findings to native image regions
copyleftdev Sep 21, 2026
8341dd3
Replay validated fleet outcomes and budget deferrals
copyleftdev Sep 21, 2026
5b3433a
Record decision evidence and partial routing timings
copyleftdev Sep 21, 2026
26f9340
Visualize routing evidence and capture a replay film draft
copyleftdev Sep 21, 2026
ad56e8d
Add discovery review policy and verify distinct reviewer routes
copyleftdev Sep 21, 2026
fcd2fc1
Capture pricing evidence and quote conservative text pilot reservations
copyleftdev Sep 21, 2026
501ded5
Prepare source-bound image evidence bundles for replay inspection
copyleftdev Sep 21, 2026
7fb0fae
Add private source-page inspector with verified OCR highlights
copyleftdev Sep 21, 2026
345db17
Verify recorded reviews against source and inspector artifacts
copyleftdev Sep 21, 2026
34e2343
Show verified findings at their recorded replay event
copyleftdev Sep 21, 2026
ac70310
Bind two-document pilot to priced routes and durable call limits
copyleftdev Sep 21, 2026
f5cf831
Enforce one-shot pilot execution with isolated credentials and durabl…
copyleftdev Sep 21, 2026
451c950
Navigate verified findings to their source image regions
copyleftdev Sep 21, 2026
f55ea66
Add private source-navigation film capture
copyleftdev Sep 21, 2026
357980f
Record source-bound media preparation metadata
copyleftdev Sep 21, 2026
0167058
Derive evidence-bound route metrics for discovery replay
copyleftdev Sep 21, 2026
7643ded
Show clock-bound route comparisons in discovery replay
copyleftdev Sep 21, 2026
b8e92ca
Package verified private replays with frozen viewer and analysis
copyleftdev Sep 21, 2026
3853da3
Gate media tests and frozen replay integrity in CI
copyleftdev Sep 21, 2026
5e2fe50
Capture frozen replay viewer hashes and route comparison
copyleftdev Sep 21, 2026
1d99b8c
Measure private image request and routing-reference boundaries
copyleftdev Sep 21, 2026
6aac27b
Add bounded image-reference execution to OpenRouter adapter
copyleftdev Sep 21, 2026
a220a3f
Preserve request-bound image receipt evidence in recordings
copyleftdev Sep 21, 2026
22028fe
Show verified image input receipts in private replay
copyleftdev Sep 21, 2026
d81b20e
Exercise recorded vision transport with real scan bundles
copyleftdev Sep 21, 2026
44eeddf
Record credential-free live pilot readiness evidence
copyleftdev Sep 21, 2026
43992f8
Connect image input receipts to submitted source pages
copyleftdev Sep 21, 2026
dd537e6
Freeze image replay packages and capture submitted-page films
copyleftdev Sep 22, 2026
81a658c
Bind supplied human assessments to recorded review evidence
copyleftdev Sep 22, 2026
90ad42b
Document funded live pilot and measured abstentions
copyleftdev Sep 22, 2026
c2ff737
Show candidate branches and policy gate outcomes in replay
copyleftdev Sep 22, 2026
ed28cb0
Add public discovery showcase and guided synthetic replay
copyleftdev Sep 22, 2026
205d1cd
Embed narrated discovery walkthrough with captions on Pages
copyleftdev Sep 22, 2026
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
26 changes: 26 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -33,6 +33,13 @@ jobs:
fi
- run: python3 scripts/test_version.py
- run: python3 scripts/test_evidence.py
- name: Install bounded image-test dependency
run: |
python3 -m venv "$RUNNER_TEMP/discovery-test-venv"
"$RUNNER_TEMP/discovery-test-venv/bin/python" -m pip install -r demo/requirements-media.txt
- name: Discovery recording and media tests
run: |
"$RUNNER_TEMP/discovery-test-venv/bin/python" -m unittest discover -s demo -p 'test_*.py'
- uses: gitleaks/gitleaks-action@e0c47f4f8be36e29cdc102c57e68cb5cbf0e8d1e # v3.0.0
env:
GITHUB_TOKEN: ${{ secrets.GITHUB_TOKEN }}
Expand All @@ -44,6 +51,20 @@ jobs:
- run: cargo fetch --locked
- run: python3 scripts/validate_local.py artifacts/ci
- run: python3 scripts/verify_evidence.py artifacts/ci --report artifacts/evidence.json
- name: Discovery recording integration
run: python3 demo/smoke.py artifacts/discovery-demo --binary target/release/braess-router
- name: Discovery reviewer integration
run: python3 demo/fleet_smoke.py artifacts/discovery-fleet --binary target/release/braess-router
- name: Verify fleet replay export
run: python3 demo/export_replay.py artifacts/discovery-fleet/run/recording artifacts/discovery-fleet/replay.json --profile fleet
- name: Discovery policy integration
run: python3 demo/fleet_smoke.py artifacts/discovery-policy --binary target/release/braess-router --discovery
- name: Verify discovery policy export
run: python3 demo/export_replay.py artifacts/discovery-policy/run/recording artifacts/discovery-policy/replay.json --profile discovery
- name: Freeze and verify synthetic discovery package
run: |
python3 demo/package_replay.py build artifacts/discovery-policy/corpus artifacts/discovery-policy/prepared/tasks.json artifacts/discovery-policy/run artifacts/discovery-policy/package
python3 demo/package_replay.py verify artifacts/discovery-policy/package
- run: cargo package --locked
- uses: actions/upload-artifact@043fb46d1a93c77aae656e7c1c64a875d1fc6a0a # v7.0.1
if: always()
Expand All @@ -52,6 +73,9 @@ jobs:
path: |
artifacts/ci
artifacts/evidence.json
artifacts/discovery-demo
artifacts/discovery-fleet
artifacts/discovery-policy
retention-days: 14
include-hidden-files: true
- uses: actions/download-artifact@3e5f45b2cfb9172054b4087a40e8e0b5a5461e7c # v8.0.1
Expand All @@ -60,3 +84,5 @@ jobs:
path: ${{ runner.temp }}/braess-evidence
- name: Verify downloaded evidence
run: python3 scripts/verify_evidence.py "$RUNNER_TEMP/braess-evidence/ci" --report "$RUNNER_TEMP/braess-evidence-report.json"
- name: Verify downloaded replay package
run: python3 demo/package_replay.py verify "$RUNNER_TEMP/braess-evidence/discovery-policy/package"
6 changes: 4 additions & 2 deletions .github/workflows/pages.yml
Original file line number Diff line number Diff line change
Expand Up @@ -2,9 +2,9 @@ name: GitHub Pages
on:
push:
branches: [main]
paths: ['site/**', 'scripts/check_site.py', 'scripts/test_site.py', '.github/workflows/pages.yml']
paths: ['site/**', 'scripts/check_site.py', 'scripts/test_site.py', 'scripts/build_discovery_site.py', 'demo/web/**', '.github/workflows/pages.yml']
pull_request:
paths: ['site/**', 'scripts/check_site.py', 'scripts/test_site.py', '.github/workflows/pages.yml']
paths: ['site/**', 'scripts/check_site.py', 'scripts/test_site.py', 'scripts/build_discovery_site.py', 'demo/web/**', '.github/workflows/pages.yml']
workflow_dispatch:
permissions:
contents: read
Expand All @@ -22,6 +22,8 @@ jobs:
- name: Validate and stage public assets
run: |
python3 scripts/test_site.py
python3 scripts/build_discovery_site.py --check
node --check site/discovery/app.js
node --check site/app.js
node --check site/traffic-data.js
python3 scripts/check_site.py --stage "$RUNNER_TEMP/braess-site"
Expand Down
548 changes: 548 additions & 0 deletions .impeccable/surface-briefs/discovery-replay.md

Large diffs are not rendered by default.

136 changes: 136 additions & 0 deletions .impeccable/surface-briefs/discovery-showcase.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,136 @@
# Public discovery showcase

Documented 2026-09-21. Mode: Experience with an Operate evidence inspector.
Scope: `site/index.html#discovery`, its narrated film embed and
`site/discovery/index.html`.
This is an ordinary extension of the precision traffic instrument. `DESIGN.md`
and `.impeccable/design.json` remain the established authority and are unchanged.

## Direction and observed surface

The public introduction explains discovery before asking visitors to interpret
routing. The landing section leads with “Discovery in motion. Follow the
decisions.” A short workflow explanation sits beside three
numbered, ruled steps: choose the review, follow the branch, inspect the record.
Its two columns use an 80px gap and collapse to one at 750px. Supporting prose
is 16px on desktop and 14px on mobile; the scope statement is 12px.

The ordinary film extension follows those steps with “Watch the walkthrough.”
The full-width, 16:9 video uses native controls, inline playback,
`preload="none"` and no autoplay. Its 74-second, 1920×1080 walkthrough has
ElevenLabs George narration and 21 default English WebVTT caption cues.
Transcript and video-download links remain available without JavaScript.
“Now follow a task yourself.” places the interactive replay action below the
film, replacing its earlier position beside the introductory steps.

Graphite rules separate the film and its next action. The heading is 24px,
supporting details 13px, and the next-action prompt 20px (18px on mobile).
At 750px the film heading, details and next-action rows stack; the video keeps
its aspect ratio. These are surface-specific treatments within the existing
visual system.

The destination begins “A document arrives. Which review next?” and defines
discovery as finding material that matters in a document collection. Three guide
columns establish the recorded experiment, explain candidate versus outcome
paths, and state what validation establishes. They become one column at 750px;
guide prose is 13px/1.75 with regular-weight 16px lead-ins. These are local
surface measurements, not additions to the system token scale.

Both surfaces retain self-hosted Archivo, near-black and porcelain, graphite
rules, square actions and the flat instrument composition. The destination
inherits the shared replay canvas, recorded-time controls, task list, route
scores, gateway timing and pale metadata pane. Dashed branches mean candidate
routes; the solid path means the recorded outcome. A crossed endpoint denotes a
preference held by the gate. Explicit labels and task outcomes repeat the visual
meaning. Each task takes one outcome path; the diagram does not imply that all
candidate handlers ran. Playback starts paused at the final frame; Start rewinds
the evidence. Spatial paths remain illustrative, independently of measured
event and gateway timestamps.

## Recording and public boundary

The published recording is the approved synthetic discovery export for run
`38d474c2-6304-47d2-86c1-2e6603731b88`: four tasks, 15 events and a final observer
time of 53,222,372ns. Standard review passes source-span validation; deeper review
returns evidence that fails validation and remains uncertain; the third task
returns local fallback; the fourth is deferred before dispatch by budget
admission. The completed count is two because local fallback is a completed
task, not a second validated review.

The router and adapter executed locally with scripted Jev and reviewer responses
and synthetic documents. This is not a semantic or legal accuracy evaluation.
Source-span validation checks evidence structure. Two synthetic generation
receipts report $0.000002; total cost remains unknown rather than reconciled
spend. The public page makes no model calls and contains no private documents.
It does not publish the private source inspector, private pilot recordings,
original source documents or credentials.

`scripts/build_discovery_site.py` owns the generated HTML, CSS and JavaScript in
`site/discovery/`, drawing the shared viewer from `demo/web/`. It adds the public
intro, search metadata and project-relative asset URLs, and removes the private
source-inspector section and its assets. Edit the shared viewer or generator and
regenerate; do not maintain a separate public viewer implementation. The
generator does not copy `replay.json`. That approved dataset is independently
pinned by SHA256 in `scripts/check_site.py`; changing it requires publication
review. Pages stages an explicit public asset allowlist, not the demo directory.

The film and its poster derive from the same reviewed public synthetic
four-task replay. The poster is an actual Chromium page capture, not a generated
image. Four new shipping assets in `site/discovery/media/`—`walkthrough.mp4`,
`poster.png`, `walkthrough.vtt` and `transcript.txt`—are explicitly allowlisted
and SHA-256 pinned. They supersede the earlier showcase's no-shipping-raster
description: the poster now ships. Raw captures, source audio, API keys,
receipts and alternate exports remain outside the publication list. Visitors
make no provider calls; the narration does not turn the scripted-provider
demonstration into a reviewer-quality claim.

## Evidence and finish

Source checked for the original showcase documentation pass: the landing HTML/CSS, public replay
HTML and recording, generator, `scripts/check_site.py`, `scripts/test_site.py`,
`scripts/test_discovery_site.cjs`, Pages workflow, and existing product, design
and discovery-replay brief. The documenter did not rerun backend execution or
browser checks.

The original showcase finish reviewer examined six intentional captures and returned
**SHIP, no material fixes required**:

- Landing section: [desktop](../review/showcase-home-desktop.png) and
[mobile](../review/showcase-home-mobile.png).
- Public introduction: [desktop](../review/showcase-intro-desktop.png) and
[mobile](../review/showcase-intro-mobile.png).
- Replay instrument: [desktop](../review/showcase-flow-desktop.png) and
[mobile](../review/showcase-flow-mobile.png).

Browser test coverage at 1440px and 390px includes landing-to-replay navigation
under the GitHub Pages project subpath, four tasks and final counts, uncertain
and deferred selections, rewind, playback, no page/HTTP errors, no horizontal
overflow and same-origin requests only. Static tests exercise stale generated
output, unauthorized recording changes, subpath assets and publication metadata.
The Pages workflow runs static tests, generator freshness and JavaScript syntax
checks before allowlist staging. For that original slice, the implementation owner reported all 10 Python
site tests, generator freshness and JavaScript syntax checks passed. Chromium
checks also passed at both sizes against 18 staged files under `/braess-router/`,
covering the interactions and request/error/overflow assertions above. That run
used the temporary precursor of the committed browser script; subsequent changes
only made the module/origin configurable and ensured the capture directory
exists. All six captures were opened by the owner and reviewer. These are
supplied execution results, not independent documenter test runs.

Small inherited diagram annotations remain a nonblocking reviewer limitation.
Detector palette/type drift was preexisting and is outside this ordinary
extension; it does not authorize a design-system refresh. That original slice
introduced no shipping raster assets; its six PNGs remain review evidence.

For the film extension, the documenter checked the landing HTML/CSS and site
README against the supplied implementation and review results. A fresh reviewer
inspected both [desktop](../review/film-embed-desktop.png) and
[mobile](../review/film-embed-mobile.png) captures and returned **SHIP, no
material fixes required**. The implementation owner reports all 11 site tests
passed, with 22 files in the public staging set. Desktop and mobile browser
checks passed native playback, no MP4 fetch before playback, no horizontal
overflow, no page or asset errors, no external calls, and navigation from the
film's replay action to all four tasks. These are supplied results, not
independent documenter test runs. The two film review captures are not shipping
assets. This brief records the reviewed local implementation; it does not
assert that a remote Pages deployment has occurred.
14 changes: 14 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
@@ -1,5 +1,19 @@
# Changelog

## 0.1.0-alpha.7

- Add explicit image-reference OpenRouter routes backed by bounded, hash-verified local PNG bundles.
- Keep routing inputs small while assembling multipart image requests after route selection.
- Retain source-reference and image hashes in generation receipts; preserve text route and journal serialization defaults.
- Direct image transport validation uses local fixtures; live model capability, pricing and semantic quality remain separate checks.

## 0.1.0-alpha.6

- Return validated routing scores, policy thresholds and partial timing traces on gateway execution responses.
- Add a provisional discovery review rubric with standard, deep and local fallback routes.
- Add a repository demo with bounded review execution, source-linked OCR findings, recorded metadata and a monochrome replay/film draft.
- Keep synthetic fixture evidence separate from live semantic quality, complete billing and publication approval.

## 0.1.0-alpha.5

- Add a text-only OpenRouter generation adapter with explicit route/model/provider mappings.
Expand Down
4 changes: 3 additions & 1 deletion Cargo.lock

Some generated files are not rendered by default. Learn more about how customized files appear on GitHub.

4 changes: 3 additions & 1 deletion Cargo.toml
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
[package]
name = "braess-router"
version = "0.1.0-alpha.5"
version = "0.1.0-alpha.7"
edition = "2024"
rust-version = "1.97.1"
publish = ["crates-io"]
Expand All @@ -13,6 +13,8 @@ include = ["src/**", "scripts/install_service.py", "eval/*.json", "eval/cases.js
default-run = "braess-router"

[dependencies]
base64 = "=0.22.1"
ring = "=0.17.14"
reqwest = { version = "0.12", default-features = false, features = ["blocking", "json", "rustls-tls"] }
serde = { version = "1", features = ["derive"] }
serde_json = "1"
Expand Down
4 changes: 3 additions & 1 deletion PRODUCT.md
Original file line number Diff line number Diff line change
Expand Up @@ -6,7 +6,7 @@
web

## Stack
Static HTML, CSS and Canvas; user accepted the recommended code-first approach. No build step or provider credentials. Suitable for static hosting, including GitHub Pages.
Static HTML, CSS and Canvas; user accepted the recommended code-first approach. No runtime build step or provider credentials. The public discovery shell is generated from the shared demo viewer. Suitable for static hosting, including GitHub Pages.

## Users
Developers evaluating Braess Router and its Jev integration.
Expand All @@ -26,5 +26,7 @@ Braess Router, black and white, elegant particle motion with meaningful shapes,
## Evidence on Hand
Verified historical gateway-durable-load-v1 experiment: 29,767 request outcomes, three phases (normal, overload, recovery). Source SHA256 manifest verifies requests and summary. Individual arrival timestamps are absent. No credentials or original request text may ship with the landing page.

The public discovery showcase adds a separately pinned, approved synthetic recording: four tasks and 15 events from local router and adapter execution with scripted providers. It shows one validated review, one uncertain review, one local fallback and one budget deferral. Candidate routes and recorded outcomes are distinct. No private documents or live model calls are included; source-span validation does not establish legal accuracy. See `.impeccable/surface-briefs/discovery-showcase.md` for surface evidence and publication boundaries.

## Product Principles
Show the mechanism. Distinguish measurements from illustration. Keep the public surface concise. Make the motion understandable without color and optional for reduced-motion users.
60 changes: 60 additions & 0 deletions demo/ADJUDICATION.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,60 @@
# Private human assessment records

Source-span validation establishes that a quote exists. It does not establish
responsiveness, privilege, completeness or legal correctness. The adjudication
companion keeps supplied human assessments separate from sealed execution logs.

```sh
python3 demo/adjudication.py prepare CORPUS PREPARED/tasks.json RUN NEW_QUEUE.json
```

The queue verifies the recording and reproduces each accepted review through
`review_link.py`. It includes every recorded task, including uncertain, deferred,
fallback and incomplete work. A task without a validated review has no review
attached. Every entry starts `awaiting_human`, with no decision. The private file
contains source quotes where available; it is not a public export.

A reviewer supplies a separate JSON file. Its `queue_sha256` is the hash of the
exact saved queue bytes, and each `review_sha256` comes from that queue entry:

```json
{
"queue_sha256": "<exact queue SHA-256>",
"reviewer_id": "<opaque reviewer identifier>",
"decisions": [
{
"task_id": "<task identifier>",
"review_sha256": "<review SHA-256, or null when absent>",
"outcome": "needs_more_context",
"note": "<reviewer's explanation>"
}
]
}
```

Record those supplied decisions with:

```sh
python3 demo/adjudication.py resolve CORPUS PREPARED/tasks.json RUN \
QUEUE.json SUPPLIED_DECISIONS.json NEW_ASSESSMENT.json
```

The resolver reconstructs the queue from its current source artifacts before
accepting anything. Outcomes are `confirm_review`, `reject_review`, or
`needs_more_context`. An absent or unvalidated review only permits the last
outcome. Duplicate/unknown tasks, mismatched review or queue hashes, modified
source reviews and oversized notes fail. Decisions may cover a subset: omitted
tasks and requests for more context remain unresolved. The report records UTC,
source hashes, supplied-decision hash, and assessed/unresolved counts. Outputs
are private, create-only and bounded to 16 MiB; source recordings stay unchanged.

This is an assessment record, not a reviewer authentication system. The opaque
identifier is supplied, not authenticated. It neither adjudicates automatically
nor converts agreement into benchmark truth. Confirmation does not approve a
redacted derivative or publication. Finding-level edits, authenticated review,
independent gold labels and UI integration remain separate work.

Local evidence: a pending queue was built for the actual OCR review recording,
and another retained all four states of the discovery-policy fixture. No real
human decision was supplied or invented. Tests use explicitly labeled synthetic
assessments to exercise resolution, tampering rejection and unresolved work.
Loading
Loading