release: v0.17.1 — the evals run for the first time, and the tails go to zero - #14
Merged
Conversation
… to zero Wave-3 of the 2026-08-29 family audit: AST-05 (first executed eval rows — haiku and sonnet, 24 blind trigger probes and 6 scored scenario runs, misses named, method stated), AST-07 (the drifted 'twenty references' aggregate dropped), AST-08 (license: MIT in all four front matters), AST-09 ($schema in both manifests, the two schemastore addresses that resolve), AST-10 (the three memory rows join the orchestrator's reference index). Board rows closed: AST-B (the Anthropic generator-evaluator citation lands in agent-evals §5a, dated) and AST-A1 (nine-subsystem coverage map — the ninth is Automation, and the suspected identity/approval-policy gap resolves as covered plus a named delegation). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Wave-3 of the 2026-08-29 family audit — the AST tails to zero, one patch release.
What's in
test/evals/RESULTS.mdgains its first two dated rows (2026-08-31, haiku + sonnet): 24 fresh blind trigger probes against the family's 28 skill descriptions and 6 scenario runs scored line by line. Both models 11/12 on triggers; scenario lines 10/12 (haiku) / 9/12 (sonnet); every miss named, Method section states the protocol and its limits.agent-evals§5a beside the doctrine it converged with.file:lineeach; the suspected identity/approval-policy gap is a split verdict — approval policy covered (governance.md), auth mechanics a named delegation (layers.md:78-80).license: MITin all four skill front matters, each re-checked withyaml.safe_load.$schemain both manifests, using the two schemastore addresses that return 200.Receipts
npm test→OK: agent-stack structurally valid (13 checks = 9 named + 4 per-skill, 4 skill(s), v0.17.1),PASS: plant_guard — 9 cases,PASS: installer — 11 case(s), residue "left nothing" on every suite.python3 test/evals_validate.py→OK: 12 trigger cases and 3 scenarios validate;--self-testgreen.claude plugin validate . --strictandplugins/agent-stack --strict→ both✔ Validation passed.🤖 Generated with Claude Code