Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
14 changes: 7 additions & 7 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -11,15 +11,15 @@ without access to their internals.**

reference architecture 4 runtimes x 2 dataset boundaries 15/15 each
independent implementation shares no code with the above 15/15
mutation analysis 17 / 17 targeted violations detected
mutation analysis 19 / 19 targeted violations detected
15 / 15 assertions independently exercised
2.2 detecting assertions per mutant (mean)
not treated as an optimization metric
defects exposed by the
conversion to vectors F-010, F-011
execution safety 0 / 576 prohibited executions
0 / 24 in the evaluation set
tests 405 passed
tests 407 passed

Authorized Recall@5 filter after truncation 0.853
filter before truncation 0.960
Expand Down Expand Up @@ -167,7 +167,7 @@ pip install agentic-dataset-conformance # contract + vectors + runner
pip install authorized-recall # the metric, no dependencies

agentic-dataset-conformance run # against its own worked example
agentic-dataset-conformance run --matrix # and the 17 broken variants
agentic-dataset-conformance run --matrix # and the 19 broken variants
agentic-dataset-conformance vectors --export ./vectors # CC0, take them
```

Expand All @@ -180,10 +180,10 @@ For the reference implementation itself, from a clone:
pip install -e ./packages/authorized-recall -e ./packages/agentic-dataset-conformance -e ".[all]"

agentic-dataset-conformance run --subject conformance.subjects:subjects # portable suite, every subject
agentic-dataset-conformance run --subject conformance.subjects:subjects --matrix # the 17-mutant detection matrix
agentic-dataset-conformance run --subject conformance.subjects:subjects --matrix # the 19-mutant detection matrix
python -m agentic_dataset.reference_suite # white-box suite, 8 configurations
python conformance/generate.py # regenerate world and vectors
pytest -q # 405 tests
pytest -q # 407 tests

python -m authorized_recall # milestone M6, the metric
python evals/evaluate.py # milestone M5, six evaluators
Expand Down Expand Up @@ -257,9 +257,9 @@ src/agentic_dataset/
packages/
authorized-recall/ the metric as its own Apache-2.0 distribution
conformance/ the normative artifact: world, vectors, verbs,
an independent implementation and 17 broken variants
an independent implementation and 19 broken variants
examples/ one runnable script per runtime, plus the MCP boundary
tests/ 405 tests
tests/ 407 tests
evals/ the M5 evaluators and the committed corpus record
docs/ architecture (three ports), results, findings, raw runs
```
Expand Down
28 changes: 28 additions & 0 deletions conformance/generate.py
Original file line number Diff line number Diff line change
Expand Up @@ -270,6 +270,34 @@ def vectors() -> dict[str, dict]:
],
}

# Two sub-properties the first fifteen vectors never exercised: a
# principal entitled to a capability above their clearance, and a grant
# used after the data it was issued for has changed.
v["ad-003-stale-grant-executes-nothing"] = {
"assertion": "AD-003",
"rules_out": "a grant for data the dataset no longer holds still opening execution",
"steps": [
req("process_engineer", expect={"decision": "GRANTED", "granted": True, "executed": True}),
{"op": "set_revision", "dataset": "purification-batches", "revision": "rev-after-grant"},
{"op": "delegate", "channel": "mcp", "dataset": "purification-batches",
"capability": "compare_batches", "scope": same,
"expect": {"executed": False, "error_contains": "revision"}},
],
}

v["ad-004-clearance-refusal-has-no-grant"] = {
"assertion": "AD-004",
"rules_out": "an entitlement above a principal's clearance minting authority",
"steps": [
{"op": "grant", "principal": "analyst", "dataset": "purification-batches",
"capability": "detect_outliers"},
req("analyst", "detect outliers in recovery", dataset="purification-batches",
capability="detect_outliers",
expect={"decision": "REFUSED", "reason": "CLASSIFICATION_EXCEEDS_CLEARANCE",
"policy_id": "AD-POL-006", "granted": False, "executed": False}),
],
}

v["ad-015-prohibited-execution-rate-zero"] = {
"assertion": "AD-015",
"rules_out": "any prohibited action executing at all, ever",
Expand Down
4 changes: 2 additions & 2 deletions docs-site/pages/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,12 @@ datasets.**

reference architecture 4 runtimes x 2 dataset boundaries 15/15 each
independent implementation shares no code with the above 15/15
mutation analysis 17 / 17 targeted violations detected
mutation analysis 19 / 19 targeted violations detected
15 / 15 assertions independently exercised
2.2 detecting assertions per mutant (mean)
execution safety 0 / 576 prohibited executions
0 / 24 in the evaluation set
tests 405 passed
tests 407 passed

Authorized Recall@5 filter after truncation 0.853
filter before truncation 0.960
Expand Down
4 changes: 2 additions & 2 deletions docs-site/src/claims.md
Original file line number Diff line number Diff line change
Expand Up @@ -12,12 +12,12 @@ measurement changes with it.
| # | Claim | Status |
|---|---|---|
| 1 | The governance model is a framework-independent behavioural contract | **Supported** |
| 2 | It can be expressed as language-neutral executable vectors | **Supported** — 15 vectors, 85 steps, JSON |
| 2 | It can be expressed as language-neutral executable vectors | **Supported** — 17 vectors, 90 steps, JSON |
| 3 | Conformance can be evaluated without access to an implementation's internals | **Supported** — the harness imports no implementation, asserted by test |
| 4 | All 15 assertions are portable | **Supported** — 15/15 through the public interface, and in **two languages** since 2026-09-03 |
| 5 | Four runtimes across two dataset boundaries all conform | **Supported** — 8 configurations, 15/15 each |
| 6 | An implementation sharing no code with the reference conforms | **Supported, with the limitation stated**: the toy is independent of the reference *code*, not of its author |
| 7 | The suite detects targeted violations | **Supported** — 17/17 mutants caught by their named assertion |
| 7 | The suite detects targeted violations | **Supported** — 19/19 mutants caught by their named assertion |
| 8 | Every assertion is exercised as the assertion under test | **Supported** — 15/15 coverage |
| 9 | The portability conversion exposed a real defect | **Strong evidence** — F-010, invisible to the white-box suite |
| 10 | The suite catches unplanned implementation mistakes, not only planted ones | **Strong evidence** — F-011, made by the toy in earnest |
Expand Down
2 changes: 1 addition & 1 deletion docs-site/src/conformance-package.md
Original file line number Diff line number Diff line change
Expand Up @@ -59,7 +59,7 @@ all fifteen.
agentic-dataset-conformance run --matrix
```

Seventeen deliberately broken variants, each removing exactly one guarantee.
Nineteen deliberately broken variants, each removing exactly one guarantee.
Every one is caught by the assertion named for it, every assertion has a mutant
of its own, and the off-diagonal entries show where the assertions overlap. A
suite that cannot fail is decoration.
Expand Down
23 changes: 22 additions & 1 deletion docs-site/src/findings.md
Original file line number Diff line number Diff line change
Expand Up @@ -147,9 +147,30 @@ run.

This is a finding about the *suite* rather than about the implementation: it is
the only direct evidence that the assertions catch a real mistake made in
earnest rather than one planted to be found. The seventeen mutants in
earnest rather than one planted to be found. The nineteen mutants in
`agentic_dataset_conformance.mutations` are planted; this one was not.

### F-012 — two checks no vector reached

Every assertion had a mutant of its own and every mutant was caught, and still
two checks every implementation makes were never exercised: refusing a
principal whose clearance is below the capability's sensitivity
(`CLASSIFICATION_EXCEEDS_CLEARANCE`), and refusing a grant used after the data
it was issued for has changed. Each refusal in the AD-004 vector is decided
before clearance is reached, and each revision change in the vectors is
followed by a fresh request, never by use of the grant issued before it. An
implementation without either check passed all fifteen vectors, and the
mutants that remove them (`clearance-ignored`, `stale-grants-accepted`) were
caught by no assertion.

Two vectors close it: `ad-004-clearance-refusal-has-no-grant` entitles a
principal to a capability above their clearance and expects the refusal, and
`ad-003-stale-grant-executes-nothing` changes the revision between a grant and
its use across the MCP boundary. With them both mutants are caught by the
assertion named for them, and every subject still conforms. The fifteen
original vectors are unchanged. Found by mutating each check an implementation
makes, not only each assertion.

---

## The semantic cache is lexical
Expand Down
4 changes: 2 additions & 2 deletions docs-site/src/index.md
Original file line number Diff line number Diff line change
Expand Up @@ -14,12 +14,12 @@ datasets.**

reference architecture 4 runtimes x 2 dataset boundaries 15/15 each
independent implementation shares no code with the above 15/15
mutation analysis 17 / 17 targeted violations detected
mutation analysis 19 / 19 targeted violations detected
15 / 15 assertions independently exercised
2.2 detecting assertions per mutant (mean)
execution safety 0 / 576 prohibited executions
0 / 24 in the evaluation set
tests 405 passed
tests 407 passed

Authorized Recall@5 filter after truncation 0.853
filter before truncation 0.960
Expand Down
6 changes: 4 additions & 2 deletions docs-site/src/portability.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,12 +123,12 @@ than "the design was correct from the start" and a more useful one.

## Mutation results

Seventeen deliberately broken variants, each removing exactly one guarantee.
Nineteen deliberately broken variants, each removing exactly one guarantee.
The matrix is what `agentic-dataset-conformance run --subject conformance.subjects:subjects --matrix` prints; the
committed run is in [`runs/mutation-matrix.txt`](https://github.com/agentic-datasets/reference/blob/main/runs/mutation-matrix.txt).

```
target detection : 17/17 mutants caught by their intended assertion
target detection : 19/19 mutants caught by their intended assertion
cross-detection : 2.2 assertions per mutant on average
coverage : 15/15 assertions have a mutant of their own
```
Expand All @@ -138,6 +138,8 @@ version of this analysis had thirteen mutants covering eleven assertions, which
meant AD-002, AD-009, AD-013 and AD-014 were exercised only as cross-detectors
— never as the assertion under test. Nothing in the pass/fail output showed
that. Drawing the matrix showed it immediately, and four mutants were added.
Full coverage of the assertions still left two checks that no vector reached;
see [`FINDINGS.md`](findings.md) F-012.

**The off-diagonal entries are a result, not noise.** They say the fifteen
assertions are not orthogonal, which is what safety invariants ought to look
Expand Down
94 changes: 54 additions & 40 deletions docs-site/src/results.md
Original file line number Diff line number Diff line change
@@ -1,17 +1,17 @@
# Results

```
15 normative assertions, 85 language-neutral vector steps
15 normative assertions, 17 vectors, 90 language-neutral vector steps

reference architecture 4 runtimes x 2 dataset boundaries 15/15 each
independent implementation shares no code with the above 15/15
mutation analysis 17 targeted violations 17/17 detected
mutation analysis 19 targeted violations 19/19 detected
15/15 assertions covered 2.2 assertions
per mutant
execution safety 0 / 39 prohibited steps, per subject
0 / 576 prohibited executions, white-box matrix
0 / 24 prohibited executions, evaluation
tests 405 passed
tests 407 passed

Authorized Recall@5 filter after truncation 0.853
filter before truncation 0.960
Expand All @@ -33,6 +33,15 @@ langgraph 1.2.11 · langchain-core 1.6.1 · llama-index-core 0.14.24
google-adk 2.8.0 · mcp 2.1.1 · pytest 9.1.1
```

The conformance matrix, the mutation analysis and the tests were measured again
on 2026-09-28, after two vectors and two mutants were added
([`FINDINGS.md`](findings.md) F-012), with:

```
langgraph 1.2.12 · langchain-core 1.6.5 · llama-index-core 0.14.25
google-adk 2.10.0 · mcp 2.2.0 · pytest 9.1.1
```

---

## 1. The conformance result
Expand Down Expand Up @@ -80,18 +89,21 @@ the same person who wrote the specification, and one person's reading of their
own document is the weakest kind of independence. The outstanding experiment is
a second reading by somebody else.

**And the suite would now notice a broken implementation.** Seventeen variants,
**And the suite would now notice a broken implementation.** Nineteen variants,
each removing exactly one guarantee, are each caught by the assertion named for
them, and every one of the fifteen assertions has a mutant of its own —
`agentic-dataset-conformance run --subject conformance.subjects:subjects --matrix`.

The detection matrix in `PORTABILITY.md` reports two separate things, because
they mean different things: **target detection** (17/17) says the suite is
they mean different things: **target detection** (19/19) says the suite is
sensitive to each named violation, and **cross-detection** (2.2 assertions per
mutant) says the assertions are not orthogonal. The second is a characterisation
rather than a score. It is also how the coverage gap was found: the first
version of the analysis had 13 mutants covering 11 assertions, and nothing in
the pass/fail output revealed that four assertions were never under test.
the pass/fail output revealed that four assertions were never under test. Two
checks were later found that no vector reached at all, clearance and a grant
used after its data changed; F-012 records them and the two vectors that close
them.

The move outside cost something, and `PORTABILITY.md` records it: AD-003 and
AD-007 became universally quantified invariants over every observation rather
Expand All @@ -106,48 +118,50 @@ argument for the matrix rather than a single run.

## 1a. Mutation analysis

Seventeen variants, each removing exactly one guarantee, run against the same
Nineteen variants, each removing exactly one guarantee, run against the same
vectors. `T` is the assertion the mutant was written for; `x` is a redundant
detection.

```
M01 M02 M03 M04 M05 M06 M07 M08 M09 M10 M11 M12 M13 M14 M15 M16 M17
--- --- --- --- --- --- --- --- --- --- --- --- --- --- --- --- ---
AD-001 T . . . x . . . . . . . . . . . . 2
AD-002 . T x . . . . . . . . . . . . . . 2
AD-003 . . T T . . . . . . . . . . . . . 2
AD-004 . . . . T . . . . . . . . . . . x 2
AD-005 . . . . . T . . . . . . . . . . . 1
AD-006 . . . . x . T . . . . . . . . . . 2
AD-007 . . . . . . . T . . . . . . x x . 3
AD-008 . . . . . . . . T T . . . . . . . 2
AD-009 . . . . . . . . . . T x x x . . x 5
AD-010 . . . . . . . . . . x T x x . . x 5
AD-011 . . . . . . . . . . . . T . . . . 1
AD-012 . . . . . . . . . . . . . T . . . 1
AD-013 . . x x . . . x . . . . . . T . . 4
AD-014 . . x x . . . x . . . . . . . T . 4
AD-015 . . . . x . . . . . . . . . . . T 2
M01 M02 M03 M04 M05 M06 M07 M08 M09 M10 M11 M12 M13 M14 M15 M16 M17 M18 M19
--- --- --- --- --- --- --- --- --- --- --- --- --- --- --- --- --- --- ---
AD-001 T . . . . x . . . . . . . . . . . . . 2
AD-002 . T x . . . . . . . . . . . . . . . . 2
AD-003 . . T T T . . . . . . . . . . . x . . 4
AD-004 . . . . . T T . . . . . . . . . . . x 3
AD-005 . . . . . . . T . . . . . . . . . . . 1
AD-006 . . . . . x . . T . . . . . . . . . . 2
AD-007 . . . . . . . . . T . . . . . . x x . 3
AD-008 . . . . . . . . . . T T . . . . . . . 2
AD-009 . . . . . . . . . . . . T x x x . . x 5
AD-010 . . . . . . . . . . . . x T x x . . x 5
AD-011 . . . . . . . . . . . . . . T . . . . 1
AD-012 . . . . . . . . . . . . . . . T . . . 1
AD-013 . . x x . . . . . x . . . . . . T . . 4
AD-014 . . x x . . . . . x . . . . . . . T . 4
AD-015 . . . . . x . . . . . . . . . . . . T 2

M01 AD-001 descriptor-not-validated
M02 AD-002 advertised-means-implemented
M03 AD-003 executes-without-a-grant
M04 AD-003 expired-tokens-accepted
M05 AD-004 refusal-still-mints-authority
M06 AD-005 indeterminate-becomes-refusal
M07 AD-006 default-allow
M08 AD-007 delegation-widens-scope
M09 AD-008 cache-ignores-principal
M10 AD-008 cache-ignores-revision
M11 AD-009 evidence-omits-principal
M12 AD-010 refusal-leaves-no-evidence
M13 AD-011 evidence-omits-revision
M14 AD-012 evidence-omits-policy-version
M15 AD-013 remote-delegation-unchecked
M16 AD-014 handoff-unchecked
M17 AD-015 prohibitions-ignored

target detection : 17/17 mutants caught by their intended assertion
M05 AD-003 stale-grants-accepted
M06 AD-004 refusal-still-mints-authority
M07 AD-004 clearance-ignored
M08 AD-005 indeterminate-becomes-refusal
M09 AD-006 default-allow
M10 AD-007 delegation-widens-scope
M11 AD-008 cache-ignores-principal
M12 AD-008 cache-ignores-revision
M13 AD-009 evidence-omits-principal
M14 AD-010 refusal-leaves-no-evidence
M15 AD-011 evidence-omits-revision
M16 AD-012 evidence-omits-policy-version
M17 AD-013 remote-delegation-unchecked
M18 AD-014 handoff-unchecked
M19 AD-015 prohibitions-ignored

target detection : 19/19 mutants caught by their intended assertion
cross-detection : 2.2 assertions per mutant on average
coverage : 15/15 assertions have a mutant of their own

Expand Down Expand Up @@ -257,7 +271,7 @@ Capability selection measured 0.800 on the first run — see
## 4. Tests

```
405 passed
407 passed
```

`pytest` parametrises the conformance suite down to one test per assertion per
Expand Down
Loading
Loading