Skip to content

release: v0.17.1 — the evals run for the first time, and the tails go to zero - #14

Merged
sshlg merged 1 commit into
mainfrom
wave3/ast-tails
Aug 31, 2026
Merged

release: v0.17.1 — the evals run for the first time, and the tails go to zero#14
sshlg merged 1 commit into
mainfrom
wave3/ast-tails

Conversation

@sshlg

@sshlg sshlg commented Aug 31, 2026

Copy link
Copy Markdown
Contributor

Wave-3 of the 2026-08-29 family audit — the AST tails to zero, one patch release.

What's in

  • AST-05test/evals/RESULTS.md gains its first two dated rows (2026-08-31, haiku + sonnet): 24 fresh blind trigger probes against the family's 28 skill descriptions and 6 scenario runs scored line by line. Both models 11/12 on triggers; scenario lines 10/12 (haiku) / 9/12 (sonnet); every miss named, Method section states the protocol and its limits.
  • AST-B — the Anthropic generator-evaluator citation (read 2026-08-30) lands in agent-evals §5a beside the doctrine it converged with.
  • AST-A1 — nine-subsystem coverage map closed on the board: the ninth subsystem is Automation; all nine map to existing doctrine with a file:line each; the suspected identity/approval-policy gap is a split verdict — approval policy covered (governance.md), auth mechanics a named delegation (layers.md:78-80).
  • AST-07 — README's drifted "twenty references" aggregate dropped (24 actual; per-skill counts stay).
  • AST-08license: MIT in all four skill front matters, each re-checked with yaml.safe_load.
  • AST-09$schema in both manifests, using the two schemastore addresses that return 200.
  • AST-10 — the three memory references join the orchestrator's index table (11 of 11).

Receipts

  • npm testOK: agent-stack structurally valid (13 checks = 9 named + 4 per-skill, 4 skill(s), v0.17.1), PASS: plant_guard — 9 cases, PASS: installer — 11 case(s), residue "left nothing" on every suite.
  • python3 test/evals_validate.pyOK: 12 trigger cases and 3 scenarios validate; --self-test green.
  • claude plugin validate . --strict and plugins/agent-stack --strict → both ✔ Validation passed.

🤖 Generated with Claude Code

… to zero

Wave-3 of the 2026-08-29 family audit: AST-05 (first executed eval rows —
haiku and sonnet, 24 blind trigger probes and 6 scored scenario runs, misses
named, method stated), AST-07 (the drifted 'twenty references' aggregate
dropped), AST-08 (license: MIT in all four front matters), AST-09 ($schema in
both manifests, the two schemastore addresses that resolve), AST-10 (the three
memory rows join the orchestrator's reference index). Board rows closed:
AST-B (the Anthropic generator-evaluator citation lands in agent-evals §5a,
dated) and AST-A1 (nine-subsystem coverage map — the ninth is Automation, and
the suspected identity/approval-policy gap resolves as covered plus a named
delegation).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@sshlg
sshlg merged commit 5ca60f1 into main Aug 31, 2026
2 checks passed
@sshlg
sshlg deleted the wave3/ast-tails branch August 31, 2026 01:40
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant