docs: [LAU-1073] describe the manual scenario run and the release evidence - #53
Conversation
…dence No shared eval harness exists. AGENT.md pointed at an internal repository and asked for 2 eval runs per release. Replace that with what a maintainer can do today: the scenario spec and fixtures in the untracked evals/ directory, a loopback mock, one Claude Code session per scenario, and the Codex bundle smoke. Name the minimum scenario subset for a release PR.
|
Updated: the eval material now lives in the separate repository sherwinski/onesignal-agent-plugin-evals. |
kalley
left a comment
There was a problem hiding this comment.
Reviewed the full diff against the repository tree. The referenced paths and interfaces all check out: evals/ is gitignored, package_directory_bundle.py has both --out and --check, scripts/checkpoint.sh honors ONESIGNAL_SKILL_ENDPOINT, and the three workflow files exist. No stale references to the old internal-repo eval process remain in AGENT.md. The new check 7 and the rewritten Releases step 2 read clearly and match the code they describe.
Clean docs change. LGTM.
Summary
No shared eval harness exists.
AGENT.mdpointed at an internal repository and asked for 2 eval runs per release. This PR replaces that text with what a maintainer can do today:evals/directory (spec, fixtures,checkpoint-transport-test.sh) and adds check 7, the manual scenario run: fixture working copy, loopback mock withGET+ 202, one Claude Code session per scenario, grade against the must and must-not columns, reset the fixture.setup_web_happy,ckpt_dirty_tree,ckpt_refusal), the Codex bundle smoke from a cleanCODEX_HOME, and a greencheckpoint-transport-test.sh. An automated runner replaces the first item when it exists.Ticket
LAU-1073
Verification
maintoday and passed; the record is on chore: Release 1.1.0 #52.bash evals/checkpoint-transport-test.sh: 195 passed, 0 failed.Note
If this merges before #52, rerun Create Release PR so
rel/1.1.0rebases onto it.