From 7cbc3cf4c21707bae01e6762d884ac4cdcd63aa2 Mon Sep 17 00:00:00 2001 From: Evgeny Kiriyak <224408464+evkir@users.noreply.github.com> Date: Tue, 22 Sep 2026 22:23:02 +0300 Subject: [PATCH 1/4] feat(register): measure what kind of assertion holds each closed row Risk 20 stood closed for fifteen days on a test that read a config field while an unscoped run spent fifty-one seconds touching a protected range. The existing guard was green and right to be: it asks whether the named test exists, not what the test asserts. Two probes settled that "measures behaviour" is not a syntactic property. A test calling validate_exploit_scope and checking its verdict is indistinguishable from one calling from_env and reading .strict_scope, so a rule failing the second accuses rows 1 and 6 wrongly. This reports instead of judging: structural, boundary, entrypoint or value, with helpers in the same module followed. Entrypoint is decided by the argument, not the receiver. CliRunner and asyncio are the receivers of every command run here; the product rides in as an argument. Measured over the register today: structural 2, boundary 2, entrypoint 6, value 6 across sixteen references. The cases are synthetic. A test reading today's register measures this morning's history and turns red when an unrelated test is renamed, which teaches the reader to edit the expectation instead of the code. The script is loaded by path, as the badge gates are, and named by a step in the typecheck job: scripts/ resolves on sys.path only because of how the package is installed here. --- .github/workflows/ci.yml | 2 + README.md | 2 +- blog/launch-post-draft.md | 2 +- scripts/register_levels.py | 203 ++++++++++++++++++ ...test_the_register_says_what_holds_a_row.py | 159 ++++++++++++++ 5 files changed, 366 insertions(+), 2 deletions(-) create mode 100644 scripts/register_levels.py create mode 100644 tests/architecture/test_the_register_says_what_holds_a_row.py diff --git a/.github/workflows/ci.yml b/.github/workflows/ci.yml index 586c9af4..270ada77 100644 --- a/.github/workflows/ci.yml +++ b/.github/workflows/ci.yml @@ -164,6 +164,8 @@ jobs: run: python scripts/typing_scope_drift.py - name: Check the declared stubs are the stubs that ship run: python scripts/stub_distributions.py + - name: Report what holds each closed row of the risk register + run: python scripts/register_levels.py mcp-1x: name: MCP surface on mcp 1.x diff --git a/README.md b/README.md index dd36211b..a7112ef6 100644 --- a/README.md +++ b/README.md @@ -5,7 +5,7 @@ ![Python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue) ![License](https://img.shields.io/badge/license-Apache_2.0-blue) ![Version](https://img.shields.io/badge/version-v1.7.0-brightgreen) -![Tests](https://img.shields.io/badge/tests-2970%20collected-brightgreen) +![Tests](https://img.shields.io/badge/tests-2979%20collected-brightgreen) ![Mypy](https://img.shields.io/badge/mypy-strict%3A%20104%2F172%20modules-blue) ![LLM](https://img.shields.io/badge/LLM-OpenAI%20%7C%20Anthropic%20%7C%20Ollama-blueviolet) ![Air-Gapped](https://img.shields.io/badge/air--gapped-ready-success) diff --git a/blog/launch-post-draft.md b/blog/launch-post-draft.md index 635fab74..efbc1a35 100644 --- a/blog/launch-post-draft.md +++ b/blog/launch-post-draft.md @@ -13,7 +13,7 @@ not exist yet, it says so. CyberAI is a multi-agent offensive-security platform: eight agents (recon, intel, exploit, report, planner, mcp-scan, redteam, web3) run a typed, audited pipeline over a shared knowledge base. -2970 tests collected under the gated selection run before every commit, with the +2979 tests collected under the gated selection run before every commit, with the slow and smoke tests deselected there and run separately, `mypy --strict` clean over 104 of 172 modules, Apache-2.0. diff --git a/scripts/register_levels.py b/scripts/register_levels.py new file mode 100644 index 00000000..99da7de9 --- /dev/null +++ b/scripts/register_levels.py @@ -0,0 +1,203 @@ +"""What kind of assertion holds each closed row of the risk register. + +Risk 20 stood closed for fifteen days on a test that read a config field +while an unscoped run spent fifty-one seconds touching a protected range. +The reference resolved, so the existing guard was green and right to be: +it asks whether the named test exists, not what the test asserts. + +Two probes on 2026-09-22 established that "measures behaviour" cannot be +decided syntactically. A test calling validate_exploit_scope and checking +its verdict is indistinguishable from one calling from_env and reading +.strict_scope; a rule that fails the second fails the first as well, and +that accusation is false. So this reports rather than judges. The register +declares what holds each row, this measures it, and the guard beside it +fails when the two disagree. A row held by a value read is not forbidden. +It is visible, which is the whole of the fix. + +Levels, in the order they are decided: + +structural the test imports nothing from cyberai; it reads the tree, + a workflow file or a document. Deliberate for some rows. +boundary something asserts about calls: assert_not_called, call_count, + call_args. The strongest form available here. +entrypoint a run-shaped call carries something from cyberai as an + argument: CliRunner().invoke(cli, ...), asyncio.run(probe()). + The receiver is the runner, never the product, so the + argument is what says the product was driven end to end. + Measured 2026-09-22: this catches the six CLI rows and + leaves the config reads alone. It does NOT yet catch a run + on an object a helper built -- rows 7, 13 and 20 read as + value for that reason. A level below the truth is not a + false claim, and the guard beside this compares what the + register declares against what this returns. +value the product is exercised and the assertions are about values + it returned or holds. + +Helper functions in the same module are followed, because a test whose +body is three calls to module helpers says nothing about itself. +""" + +from __future__ import annotations + +import ast +import collections +import pathlib +import re +import sys + +_ROOT = pathlib.Path(__file__).resolve().parents[1] +_REGISTER = _ROOT / "docs" / "architecture" / "risk-register.md" + +_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|(.+)\|\s*$") +_REFERENCE = re.compile(r"tests/[A-Za-z0-9_/]+\.py::[A-Za-z0-9_]+") + +_CALL_ASSERTION = re.compile(r"assert_(?:not_)?(?:called|awaited)\w*|call_count|call_args") +_RUN_NAMES = frozenset({"invoke", "run"}) + +LEVELS = ("structural", "boundary", "entrypoint", "value") + + +def product_names(tree: ast.Module) -> set[str]: + """Names this module pulled out of cyberai, under whatever alias.""" + names: set[str] = set() + for node in ast.walk(tree): + if isinstance(node, ast.ImportFrom): + if (node.module or "").startswith("cyberai"): + names |= {alias.asname or alias.name for alias in node.names} + elif isinstance(node, ast.Import): + for alias in node.names: + if alias.name.startswith("cyberai"): + names.add(alias.asname or alias.name.split(".")[0]) + return names + + +def _root_name(node: ast.expr) -> str | None: + while isinstance(node, (ast.Attribute, ast.Subscript, ast.Call)): + node = node.value if isinstance(node, (ast.Attribute, ast.Subscript)) else node.func + return node.id if isinstance(node, ast.Name) else None + + +def _argument_names(call: ast.Call) -> set[str]: + """Every name and attribute appearing in the arguments of one call.""" + out: set[str] = set() + for argument in [*call.args, *(keyword.value for keyword in call.keywords)]: + for child in ast.walk(argument): + if isinstance(child, ast.Name): + out.add(child.id) + elif isinstance(child, ast.Attribute): + out.add(child.attr) + return out + + +def _mentioned(node: ast.AST) -> set[str]: + out: set[str] = set() + for child in ast.walk(node): + if isinstance(child, ast.Name): + out.add(child.id) + elif isinstance(child, ast.Attribute): + out.add(child.attr) + return out + + +def _functions(tree: ast.Module) -> dict[str, ast.FunctionDef | ast.AsyncFunctionDef]: + return { + node.name: node + for node in ast.walk(tree) + if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef)) + } + + +def level_of(source: str, function: str) -> str: + """The level of one test, following helpers defined in the same module.""" + tree = ast.parse(source) + product = product_names(tree) + functions = _functions(tree) + if function not in functions: + return "structural" + + seen: set[str] = set() + queue: collections.deque[str] = collections.deque([function]) + reaches = False + boundary = False + entrypoint = False + + while queue: + current = queue.popleft() + if current in seen: + continue + seen.add(current) + node = functions.get(current) + if node is None: + continue + segment = ast.get_source_segment(source, node) or "" + if _CALL_ASSERTION.search(segment): + boundary = True + for call in ast.walk(node): + if not isinstance(call, ast.Call): + continue + if isinstance(call.func, ast.Name): + called = call.func.id + elif isinstance(call.func, ast.Attribute): + called = call.func.attr + root = _root_name(call.func.value) + if root is not None and root in product: + reaches = True + else: + continue + if called in product: + reaches = True + if called in _RUN_NAMES and _argument_names(call) & product: + entrypoint = True + reaches = True + if product & _mentioned(node): + reaches = True + for name in _mentioned(node): + if name in functions and name not in seen: + queue.append(name) + + if not reaches: + return "structural" + if boundary: + return "boundary" + if entrypoint: + return "entrypoint" + return "value" + + +def closed_references(text: str) -> list[tuple[int, str]]: + """(row number, node id) for every reference a closed row carries.""" + out: list[tuple[int, str]] = [] + for line in text.splitlines(): + match = _ROW.match(line) + if match and match.group(3) == "closed": + for reference in _REFERENCE.findall(match.group(4)): + out.append((int(match.group(1)), reference)) + return out + + +def measure(root: pathlib.Path, text: str) -> list[tuple[int, str, str]]: + """(row, reference, level) for every closed reference in the register.""" + out: list[tuple[int, str, str]] = [] + for number, reference in closed_references(text): + relative, _, function = reference.partition("::") + path = root / relative + if not path.exists(): + out.append((number, reference, "structural")) + continue + out.append((number, reference, level_of(path.read_text(encoding="utf-8"), function))) + return out + + +def main() -> int: + measured = measure(_ROOT, _REGISTER.read_text(encoding="utf-8")) + counts: collections.Counter[str] = collections.Counter(level for _, _, level in measured) + for number, reference, level in measured: + print(f"{number:>3} {level:<11} {reference}") + print() + print(" ".join(f"{level}: {counts[level]}" for level in LEVELS)) + print(f"references: {len(measured)}") + return 0 + + +if __name__ == "__main__": + sys.exit(main()) diff --git a/tests/architecture/test_the_register_says_what_holds_a_row.py b/tests/architecture/test_the_register_says_what_holds_a_row.py new file mode 100644 index 00000000..39c52423 --- /dev/null +++ b/tests/architecture/test_the_register_says_what_holds_a_row.py @@ -0,0 +1,159 @@ +"""The level of a closed row is measured, not asserted by its author. + +Risk 20 stood closed for fifteen days on a test that read a config field. +The reference resolved, so `test_the_register_names_tests_that_exist` +was green and correct to be: it asks whether the named test exists. + +Two probes on 2026-09-22 settled that "measures behaviour" is not a +syntactic property. A test calling `validate_exploit_scope` and checking +its verdict looks exactly like one calling `from_env` and reading +`.strict_scope`; a rule failing the second fails the first too, and rows +1 and 6 would be accused wrongly. So nothing here forbids a level. The +register states what holds each row, `scripts/register_levels.py` +measures it, and the guard fails when the two disagree -- a row held by +a value read is allowed, and is now readable as such from the page. + +The cases below are written here rather than taken from the tree. A test +that reads today's register measures this morning's history and turns red +whenever an unrelated test is renamed, which teaches the reader to edit +the expectation. These fix the classifier instead: each is a shape the +measurement must keep telling apart. +""" + +import importlib.util +import pathlib +import types + +_ROOT = pathlib.Path(__file__).resolve().parents[2] +_SCRIPT = _ROOT / "scripts" / "register_levels.py" + + +def _levels_tool() -> types.ModuleType: + """Loaded by path, as the badge gates are. + + Putting scripts/ on sys.path works only because of how the package is + installed here, and a gate resting on the install mode is a gate about + one machine. + """ + spec = importlib.util.spec_from_file_location("register_levels", _SCRIPT) + assert spec and spec.loader, f"no module at {_SCRIPT}" + module = importlib.util.module_from_spec(spec) + spec.loader.exec_module(module) + return module + + +register_levels = _levels_tool() + +_CLI_RUN = """ +from click.testing import CliRunner +from cyberai.cli.mcp_scan import mcp_scan + + +def _run(): + return CliRunner().invoke(mcp_scan, ["http://t"]) + + +def test_it(): + assert _run().exit_code == 0 +""" + +_VALUE_READ = """ +from cyberai.core.config import CyberAIConfig + + +def test_it(monkeypatch): + monkeypatch.setenv("CYBERAI_STRICT_SCOPE", "") + assert CyberAIConfig.from_env().strict_scope is True +""" + +_CALL_ASSERTION = """ +from unittest.mock import patch +from cyberai.core.orchestrator import Orchestrator + + +def test_it(): + with patch("cyberai.agents.recon.agent.ReconAgent.run") as probe: + Orchestrator().run("10.0.0.1") + probe.assert_not_called() +""" + +_RUNNER_WITHOUT_US = """ +from click.testing import CliRunner +from other.pkg import their_cli + +from cyberai.core.config import CyberAIConfig + + +def test_it(): + CliRunner().invoke(their_cli, []) + assert CyberAIConfig().strict_scope is True +""" + +_UNUSED_IMPORT = """ +from cyberai.core.config import CyberAIConfig + + +def test_it(): + assert 1 == 1 +""" + +_READS_THE_TREE = """ +import pathlib + + +def test_it(): + assert "on:" in pathlib.Path(".github/workflows/ci.yml").read_text() +""" + + +def test_a_command_driven_through_the_runner_is_an_entrypoint() -> None: + """The receiver is the runner; the product rides in as an argument.""" + assert register_levels.level_of(_CLI_RUN, "test_it") == "entrypoint" + + +def test_a_field_read_back_from_the_product_is_a_value() -> None: + """The shape risk 20 wore while an unscoped run touched a protected range.""" + assert register_levels.level_of(_VALUE_READ, "test_it") == "value" + + +def test_an_assertion_about_calls_outranks_the_run_around_it() -> None: + """Both are present here; the stronger claim is the one reported.""" + assert register_levels.level_of(_CALL_ASSERTION, "test_it") == "boundary" + + +def test_a_test_that_imports_nothing_from_the_product_is_structural() -> None: + assert register_levels.level_of(_READS_THE_TREE, "test_it") == "structural" + + +def test_a_helper_in_the_same_module_is_followed() -> None: + """A body of three helper calls says nothing about itself.""" + assert register_levels.level_of(_CLI_RUN, "_run") == "entrypoint" + + +def test_a_runner_carrying_nothing_of_ours_is_not_an_entrypoint() -> None: + """Otherwise every CliRunner in the tree reads as a product run. + + The product is imported and read here, so the test reaches it: what is + missing is the product going into the runner. Drop the import as well + and the answer is structural, which the case below states separately. + """ + assert register_levels.level_of(_RUNNER_WITHOUT_US, "test_it") == "value" + + +def test_importing_the_product_without_touching_it_is_structural() -> None: + """An unused import is not a claim about behaviour.""" + assert register_levels.level_of(_UNUSED_IMPORT, "test_it") == "structural" + + +def test_a_name_the_register_does_not_carry_is_not_invented() -> None: + assert register_levels.level_of(_CLI_RUN, "test_absent") == "structural" + + +def test_the_measurement_covers_every_closed_reference() -> None: + """A classifier that silently skips rows would report a clean page.""" + text = (_ROOT / "docs" / "architecture" / "risk-register.md").read_text(encoding="utf-8") + carried = len(register_levels.closed_references(text)) + measured = register_levels.measure(_ROOT, text) + assert carried > 0, "no closed row carries a reference -- the row pattern broke" + assert len(measured) == carried + assert {level for _, _, level in measured} <= set(register_levels.LEVELS) From 81c6b0fe0b68b59d578e86ce698f8f30d6740095 Mon Sep 17 00:00:00 2001 From: Evgeny Kiriyak <224408464+evkir@users.noreply.github.com> Date: Tue, 22 Sep 2026 22:33:24 +0300 Subject: [PATCH 2/4] docs(register): every closed row says what holds it, and the tree checks The page now carries a Held by column, written by the script rather than by whoever edits the row, and a guard that fails when the column and the tree disagree. Measured today: structural 2, boundary 2, entrypoint 6, value 6 over sixteen references. Risk 20 reads boundary/value, which is what it would have said out loud while it stood closed on a config read. No level is forbidden. Two probes established that the distinction cannot be made from syntax, so a value row is not a defect: it is a row whose guard could be stronger, now readable as such without opening the test. Both readers of this page matched the new column into the evidence cell and went on working. That is the wrong reason for a check to be green, so each pattern was widened to know the shape it reads. A column of claims decays the way the last table did: typing-scope.md carried a module at 17 errors while the tree said 12, and nobody was wrong when it was written. Hence the comparison rather than a note. Known gap, recorded rather than fixed: a row the pattern cannot parse leaves both sides of that comparison at once and is not reported missing. Counting the rows is a second guard, not this one. --- README.md | 2 +- blog/launch-post-draft.md | 2 +- docs/architecture/risk-register.md | 86 ++++++++++++------- scripts/register_levels.py | 22 ++++- ...est_the_register_names_tests_that_exist.py | 12 ++- ...test_the_register_says_what_holds_a_row.py | 37 ++++++++ 6 files changed, 122 insertions(+), 39 deletions(-) diff --git a/README.md b/README.md index a7112ef6..99f3003b 100644 --- a/README.md +++ b/README.md @@ -5,7 +5,7 @@ ![Python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue) ![License](https://img.shields.io/badge/license-Apache_2.0-blue) ![Version](https://img.shields.io/badge/version-v1.7.0-brightgreen) -![Tests](https://img.shields.io/badge/tests-2979%20collected-brightgreen) +![Tests](https://img.shields.io/badge/tests-2981%20collected-brightgreen) ![Mypy](https://img.shields.io/badge/mypy-strict%3A%20104%2F172%20modules-blue) ![LLM](https://img.shields.io/badge/LLM-OpenAI%20%7C%20Anthropic%20%7C%20Ollama-blueviolet) ![Air-Gapped](https://img.shields.io/badge/air--gapped-ready-success) diff --git a/blog/launch-post-draft.md b/blog/launch-post-draft.md index efbc1a35..aa7c5e9a 100644 --- a/blog/launch-post-draft.md +++ b/blog/launch-post-draft.md @@ -13,7 +13,7 @@ not exist yet, it says so. CyberAI is a multi-agent offensive-security platform: eight agents (recon, intel, exploit, report, planner, mcp-scan, redteam, web3) run a typed, audited pipeline over a shared knowledge base. -2979 tests collected under the gated selection run before every commit, with the +2981 tests collected under the gated selection run before every commit, with the slow and smoke tests deselected there and run separately, `mypy --strict` clean over 104 of 172 modules, Apache-2.0. diff --git a/docs/architecture/risk-register.md b/docs/architecture/risk-register.md index e2da208b..d5fe168e 100644 --- a/docs/architecture/risk-register.md +++ b/docs/architecture/risk-register.md @@ -20,46 +20,68 @@ Status vocabulary: `closed` — a test in this tree fails if the risk returns. `partly` — one half is guarded, the other is named below. `open` — measured, not fixed. `unguarded` — held by discipline, with no machine behind it. +Held by says what kind of assertion the named test makes, and it is written +by `scripts/register_levels.py` rather than by whoever edits the row. Risk 20 +stood closed for fifteen days on a test that read a config field while an +unscoped run spent fifty-one seconds touching a protected range; the +reference resolved the whole time, so the guard over this page was green and +right to be. This column is what that row would have said out loud. + +`boundary` — something is asserted about calls: a probe that was not reached, +a count, the arguments it was given. `entrypoint` — a command or agent of +ours was driven end to end and the result inspected. `value` — the product +was exercised and the assertions are about what it returned or holds. +`structural` — the test reads the tree, a workflow file or a page, and +touches no product code; deliberate for rows about the repository itself. +A `-` means the row is not closed, so there is nothing to measure. + +No level is forbidden here. Two probes on 2026-09-22 established that +"measures behaviour" cannot be decided from syntax: a test calling +`validate_exploit_scope` and checking its verdict looks exactly like one +calling `from_env` and reading `.strict_scope`, and a rule failing the second +would accuse rows 1 and 6 wrongly. A `value` row is not a defect. It is a row +whose guard could be stronger, now readable as such without opening the test. + ## Shipped defects, found by live measurement -| # | Risk | Status | Measured by | -|---|---|---|---| -| 1 | Web exploitation unreachable under every `--scope` value: two gates normalised opposite sides of one comparison | closed | `tests/unit/test_scope_guard_normalises_the_target.py::test_both_gates_reach_the_same_verdict` | -| 2 | A skipped phase printed as success, exit code 0 on `state: failed` | closed | `tests/unit/test_a_scope_refusal_is_visible_at_the_end.py::test_the_status_line_reports_partial_for_a_refused_walk` | -| 3 | Tests feed argument shapes production never sends | closed | `tests/architecture/test_a_url_parameter_is_fed_a_url.py::test_no_test_feeds_a_url_parameter_something_without_a_scheme` | -| 4 | The published score came from a surface profile blind on both live targets | closed | `tests/unit/test_the_manifest_records_the_surface_it_saw.py::test_the_manifest_carries_the_profile` | -| 5 | An AI-native benchmark measured a path that called no model | closed | `tests/unit/test_the_bench_table_says_who_solved_it.py::test_a_model_free_run_says_so_next_to_the_score` | -| 6 | Documented out-of-band cost understated sixteenfold | closed | `tests/unit/test_web_exploit_oob.py::test_every_delivery_is_charged_to_the_budget` | -| 7 | One of four web3 engines decided the verdict; aderyn returned nothing | closed | `tests/unit/test_web3_agent_merge.py::test_run_merges_and_cross_validates` | +| # | Risk | Status | Held by | Measured by | +|---|---|---|---|---| +| 1 | Web exploitation unreachable under every `--scope` value: two gates normalised opposite sides of one comparison | closed | value | `tests/unit/test_scope_guard_normalises_the_target.py::test_both_gates_reach_the_same_verdict` | +| 2 | A skipped phase printed as success, exit code 0 on `state: failed` | closed | entrypoint | `tests/unit/test_a_scope_refusal_is_visible_at_the_end.py::test_the_status_line_reports_partial_for_a_refused_walk` | +| 3 | Tests feed argument shapes production never sends | closed | structural | `tests/architecture/test_a_url_parameter_is_fed_a_url.py::test_no_test_feeds_a_url_parameter_something_without_a_scheme` | +| 4 | The published score came from a surface profile blind on both live targets | closed | entrypoint | `tests/unit/test_the_manifest_records_the_surface_it_saw.py::test_the_manifest_carries_the_profile` | +| 5 | An AI-native benchmark measured a path that called no model | closed | entrypoint | `tests/unit/test_the_bench_table_says_who_solved_it.py::test_a_model_free_run_says_so_next_to_the_score` | +| 6 | Documented out-of-band cost understated sixteenfold | closed | value | `tests/unit/test_web_exploit_oob.py::test_every_delivery_is_charged_to_the_budget` | +| 7 | One of four web3 engines decided the verdict; aderyn returned nothing | closed | value | `tests/unit/test_web3_agent_merge.py::test_run_merges_and_cross_validates` | ## Technical -| # | Risk | Status | Measured by | -|---|---|---|---| -| 8 | The CVE-Bench adapter was written against criteria older than v2.1.0 | partly | `tests/integration/test_the_bench_answers_to_the_upstream.py::test_the_adapter_answers_to_the_checkout_on_disk` — carries the smoke marker, so CI skips it and only a workstation with the checkout runs it | -| 9 | The MCP client did not report which protocol revision it negotiated | closed | `tests/integration/test_mcp_revision_in_output.py::test_the_terminal_names_the_revision_the_probe_negotiated` | -| 10 | The detector does not cover tool icons or URL-mode elicitation, both added to the protocol in revision 2025-11-25 | partly | Icons: `tests/unit/test_mcp_poisoning.py::test_a_directive_in_an_icon_field_reaches_the_matcher` and `tests/unit/test_mcp_poisoning.py::test_an_executable_icon_carrier_is_a_signal`. The field was outside the collected whitelist, so no pattern could have reached it. Elicitation: the risk was written on a premise the SDK does not support -- `ServerCapabilities` has no elicitation field, so a scanned server cannot advertise URL mode and a probe that calls nothing never receives error -32042. The exposure runs the other way and is held by `tests/unit/test_the_probe_does_not_offer_to_open_a_url.py::test_the_probe_advertises_no_elicitation`. Partly: what a target does with elicitation is reachable only by calling its tools, which this scanner does not do | -| 11 | Detection and false-positive figures come from a corpus this project wrote | open | `tests/architecture/test_corpus_integrity.py::test_both_classes_meet_the_floor` guards the corpus, not its provenance. No external corpus has ever been run | -| 12 | CVE-Bench and phantom-grid both claim port 9090 | closed | `tests/unit/test_cve_bench_driver.py::test_a_taken_port_is_named_not_blamed_on_the_stack` | -| 13 | Two finding counts in one result, neither reconciled with the other | closed | `tests/unit/test_web3_agent_merge.py::test_aderyn_only_critical_raises_the_headline` | -| 14 | Access findings and escalation paths were computed and never printed | closed | `tests/integration/test_web3_audit_cli.py::test_audit_prints_every_bucket_not_just_slither` | +| # | Risk | Status | Held by | Measured by | +|---|---|---|---|---| +| 8 | The CVE-Bench adapter was written against criteria older than v2.1.0 | partly | - | `tests/integration/test_the_bench_answers_to_the_upstream.py::test_the_adapter_answers_to_the_checkout_on_disk` — carries the smoke marker, so CI skips it and only a workstation with the checkout runs it | +| 9 | The MCP client did not report which protocol revision it negotiated | closed | entrypoint | `tests/integration/test_mcp_revision_in_output.py::test_the_terminal_names_the_revision_the_probe_negotiated` | +| 10 | The detector does not cover tool icons or URL-mode elicitation, both added to the protocol in revision 2025-11-25 | partly | - | Icons: `tests/unit/test_mcp_poisoning.py::test_a_directive_in_an_icon_field_reaches_the_matcher` and `tests/unit/test_mcp_poisoning.py::test_an_executable_icon_carrier_is_a_signal`. The field was outside the collected whitelist, so no pattern could have reached it. Elicitation: the risk was written on a premise the SDK does not support -- `ServerCapabilities` has no elicitation field, so a scanned server cannot advertise URL mode and a probe that calls nothing never receives error -32042. The exposure runs the other way and is held by `tests/unit/test_the_probe_does_not_offer_to_open_a_url.py::test_the_probe_advertises_no_elicitation`. Partly: what a target does with elicitation is reachable only by calling its tools, which this scanner does not do | +| 11 | Detection and false-positive figures come from a corpus this project wrote | open | - | `tests/architecture/test_corpus_integrity.py::test_both_classes_meet_the_floor` guards the corpus, not its provenance. No external corpus has ever been run | +| 12 | CVE-Bench and phantom-grid both claim port 9090 | closed | boundary | `tests/unit/test_cve_bench_driver.py::test_a_taken_port_is_named_not_blamed_on_the_stack` | +| 13 | Two finding counts in one result, neither reconciled with the other | closed | value | `tests/unit/test_web3_agent_merge.py::test_aderyn_only_critical_raises_the_headline` | +| 14 | Access findings and escalation paths were computed and never printed | closed | entrypoint | `tests/integration/test_web3_audit_cli.py::test_audit_prints_every_bucket_not_just_slither` | ## Product and market -| # | Risk | Status | Measured by | -|---|---|---|---| -| 15 | A crowded category: the metadata scanner people know was acquired, and a large vendor gives a pattern scanner away | open | not a code property. See `docs/competitive-landscape-2026.md`, collected 2026-08-11; re-collect rather than trust it | -| 16 | The value of an agent is a low false-positive rate, not a solve rate, and agent false positives are measured above human ones | partly | `tests/architecture/test_detector_baseline.py::test_whole_subclasses_are_invisible_today` names what is missed. The rate itself is measured only on our corpus — see 11 | -| 17 | A web3 report is not a submission: proof of concept is required at every severity | open | the working Foundry exploit lives outside the repository, so nothing runs it | -| 18 | A benchmark nobody outside can reproduce | partly | `tests/integration/test_repro_chain.py::test_two_identical_runs_share_fingerprint` pins seed, config and suite. Target image digests and external tool versions are not recorded | +| # | Risk | Status | Held by | Measured by | +|---|---|---|---|---| +| 15 | A crowded category: the metadata scanner people know was acquired, and a large vendor gives a pattern scanner away | open | - | not a code property. See `docs/competitive-landscape-2026.md`, collected 2026-08-11; re-collect rather than trust it | +| 16 | The value of an agent is a low false-positive rate, not a solve rate, and agent false positives are measured above human ones | partly | - | `tests/architecture/test_detector_baseline.py::test_whole_subclasses_are_invisible_today` names what is missed. The rate itself is measured only on our corpus — see 11 | +| 17 | A web3 report is not a submission: proof of concept is required at every severity | open | - | the working Foundry exploit lives outside the repository, so nothing runs it | +| 18 | A benchmark nobody outside can reproduce | partly | - | `tests/integration/test_repro_chain.py::test_two_identical_runs_share_fingerprint` pins seed, config and suite. Target image digests and external tool versions are not recorded | ## Legal, reputational, structural -| # | Risk | Status | Measured by | -|---|---|---|---| -| 19 | Publishing numbers produced by a path that was broken | closed | `tests/architecture/test_nothing_publishes_without_the_checks.py::test_the_upload_cannot_run_before_the_guard` | -| 20 | An offensive tool that does not require an authorisation scope | closed | `tests/unit/test_the_refusal_comes_before_the_first_packet.py::test_an_unscoped_run_never_reaches_the_recon_agent` holds the refusal ahead of the pipeline: measured on 2026-09-17, an unscoped run against a protected range had already spent 51 seconds on nmap, whois, dns and subdomain enumeration before the exploit phase declined it. The row said closed while the target was being touched, because the test behind it read a config field rather than the network. `--no-strict-scope` and `CYBERAI_STRICT_SCOPE=0` remain the named ways to proceed anyway, and since 2026-09-19 they are the only ones: `tests/unit/test_strict_scope.py::test_only_a_named_word_turns_the_refusal_off` holds the refusal against a value nobody chose. The flag reader answers "is this one of the words for yes", so an empty variable -- what a shell leaves for `VAR=` in a .env file -- and a misspelling both read as off, which disarmed the one flag here whose default protects the run | -| 21 | Private plans and journals reaching a public repository | unguarded | no test and no workflow step checks this. A tree-wide history search finds none of those filenames, which is a measurement of the past, not a guard on the next commit | -| 22 | One developer, measuring by hand | open | structural. The mitigation is that every closed row above names a test rather than a memory | -| 23 | A measurement that cannot tell a quiet environment from a clean tree | partly | `tests/architecture/test_the_drift_report_refuses_an_unmeasured_environment.py::test_a_checker_that_never_ran_is_not_a_clean_package` and the four rows beside it. Measured on 2026-09-18 on an untouched checkout: with no mypy installed the drift report printed 172 clean modules, 0 errors and no drift and exited zero; without `types-networkx` it printed 275 errors and accused `cyberai/core/kb_graph.py` of being undeclared; on mcp 1.28.1 it printed 284 instead of 285. Five signals, one tree, every guard green, because the guards read declarations rather than the environment. The report now refuses a run it cannot vouch for and names the package that differs. Partly: the same question is unanswered for the test suite, which collected nothing on a machine without `pytest-asyncio` and said so as three collection errors rather than as an unmeasured environment | -| 24 | Configuration naming a provider that does not exist | closed | `tests/unit/test_the_config_cannot_name_an_unknown_provider.py::test_an_unknown_provider_does_not_reach_the_config` for the environment and `tests/unit/test_the_cli_cannot_name_an_unknown_provider.py::test_an_unknown_provider_is_refused_by_the_parser` for the flag. `CYBERAI_LLM_PROVIDER=gemini` used to produce a config whose provider was that string, and `api_key_for` then resolved a credential by name: a run asked for one vendor could be sent to another under a key nobody named. Both readers narrow against the declared Literal, not a second list -- `test_the_reader_reads_the_declared_literal_and_not_a_copy` fails if a copy appears. The two boundaries answer differently on purpose: the environment falls back to the default because a stale variable must not abort a scan, the flag is refused because whoever typed it is there to read the reply. The three files carrying this path -- and the validator beside them -- entered `[tool.mypy] files` in the same commits, which is what makes the assignments visible to CI at all | +| # | Risk | Status | Held by | Measured by | +|---|---|---|---|---| +| 19 | Publishing numbers produced by a path that was broken | closed | structural | `tests/architecture/test_nothing_publishes_without_the_checks.py::test_the_upload_cannot_run_before_the_guard` | +| 20 | An offensive tool that does not require an authorisation scope | closed | boundary/value | `tests/unit/test_the_refusal_comes_before_the_first_packet.py::test_an_unscoped_run_never_reaches_the_recon_agent` holds the refusal ahead of the pipeline: measured on 2026-09-17, an unscoped run against a protected range had already spent 51 seconds on nmap, whois, dns and subdomain enumeration before the exploit phase declined it. The row said closed while the target was being touched, because the test behind it read a config field rather than the network. `--no-strict-scope` and `CYBERAI_STRICT_SCOPE=0` remain the named ways to proceed anyway, and since 2026-09-19 they are the only ones: `tests/unit/test_strict_scope.py::test_only_a_named_word_turns_the_refusal_off` holds the refusal against a value nobody chose. The flag reader answers "is this one of the words for yes", so an empty variable -- what a shell leaves for `VAR=` in a .env file -- and a misspelling both read as off, which disarmed the one flag here whose default protects the run | +| 21 | Private plans and journals reaching a public repository | unguarded | - | no test and no workflow step checks this. A tree-wide history search finds none of those filenames, which is a measurement of the past, not a guard on the next commit | +| 22 | One developer, measuring by hand | open | - | structural. The mitigation is that every closed row above names a test rather than a memory | +| 23 | A measurement that cannot tell a quiet environment from a clean tree | partly | - | `tests/architecture/test_the_drift_report_refuses_an_unmeasured_environment.py::test_a_checker_that_never_ran_is_not_a_clean_package` and the four rows beside it. Measured on 2026-09-18 on an untouched checkout: with no mypy installed the drift report printed 172 clean modules, 0 errors and no drift and exited zero; without `types-networkx` it printed 275 errors and accused `cyberai/core/kb_graph.py` of being undeclared; on mcp 1.28.1 it printed 284 instead of 285. Five signals, one tree, every guard green, because the guards read declarations rather than the environment. The report now refuses a run it cannot vouch for and names the package that differs. Partly: the same question is unanswered for the test suite, which collected nothing on a machine without `pytest-asyncio` and said so as three collection errors rather than as an unmeasured environment | +| 24 | Configuration naming a provider that does not exist | closed | entrypoint/value | `tests/unit/test_the_config_cannot_name_an_unknown_provider.py::test_an_unknown_provider_does_not_reach_the_config` for the environment and `tests/unit/test_the_cli_cannot_name_an_unknown_provider.py::test_an_unknown_provider_is_refused_by_the_parser` for the flag. `CYBERAI_LLM_PROVIDER=gemini` used to produce a config whose provider was that string, and `api_key_for` then resolved a credential by name: a run asked for one vendor could be sent to another under a key nobody named. Both readers narrow against the declared Literal, not a second list -- `test_the_reader_reads_the_declared_literal_and_not_a_copy` fails if a copy appears. The two boundaries answer differently on purpose: the environment falls back to the default because a stale variable must not abort a scan, the flag is refused because whoever typed it is there to read the reply. The three files carrying this path -- and the validator beside them -- entered `[tool.mypy] files` in the same commits, which is what makes the assignments visible to CI at all | diff --git a/scripts/register_levels.py b/scripts/register_levels.py index 99da7de9..94a3bd97 100644 --- a/scripts/register_levels.py +++ b/scripts/register_levels.py @@ -48,7 +48,7 @@ _ROOT = pathlib.Path(__file__).resolve().parents[1] _REGISTER = _ROOT / "docs" / "architecture" / "risk-register.md" -_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|(.+)\|\s*$") +_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|\s*([\w/-]+)\s*\|(.+)\|\s*$") _REFERENCE = re.compile(r"tests/[A-Za-z0-9_/]+\.py::[A-Za-z0-9_]+") _CALL_ASSERTION = re.compile(r"assert_(?:not_)?(?:called|awaited)\w*|call_count|call_args") @@ -170,11 +170,29 @@ def closed_references(text: str) -> list[tuple[int, str]]: for line in text.splitlines(): match = _ROW.match(line) if match and match.group(3) == "closed": - for reference in _REFERENCE.findall(match.group(4)): + for reference in _REFERENCE.findall(match.group(5)): out.append((int(match.group(1)), reference)) return out +def rows(text: str) -> list[tuple[int, str, str]]: + """(number, status, declared level) for each numbered row.""" + out: list[tuple[int, str, str]] = [] + for line in text.splitlines(): + match = _ROW.match(line) + if match: + out.append((int(match.group(1)), match.group(3), match.group(4))) + return out + + +def declared_level(text: str, number: int) -> str: + """What the page says holds one row, or the empty string if it says nothing.""" + for found, _, level in rows(text): + if found == number: + return level + return "" + + def measure(root: pathlib.Path, text: str) -> list[tuple[int, str, str]]: """(row, reference, level) for every closed reference in the register.""" out: list[tuple[int, str, str]] = [] diff --git a/tests/architecture/test_the_register_names_tests_that_exist.py b/tests/architecture/test_the_register_names_tests_that_exist.py index 609fa248..59b370b6 100644 --- a/tests/architecture/test_the_register_names_tests_that_exist.py +++ b/tests/architecture/test_the_register_names_tests_that_exist.py @@ -26,17 +26,23 @@ _REGISTER = _ROOT / "docs" / "architecture" / "risk-register.md" _REFERENCE = re.compile(r"tests/[A-Za-z0-9_/]+\.py::[A-Za-z0-9_]+") -_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|(.+?)\|\s*$") +_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|\s*([\w/-]+)\s*\|(.+?)\|\s*$") _DECLARED = {"closed", "partly", "open", "unguarded"} def rows(text: str) -> list[tuple[int, str, str]]: - """(number, status, evidence cell) for each numbered row in the register.""" + """(number, status, evidence cell) for each numbered row in the register. + + The table grew a fourth column on 2026-09-22 saying what kind of + assertion holds each closed row. Both readers of this page matched it + into the evidence cell and kept working, which is the wrong reason for + a check to be green: the pattern has to know the shape it reads. + """ out = [] for line in text.splitlines(): match = _ROW.match(line) if match: - out.append((int(match.group(1)), match.group(3), match.group(4))) + out.append((int(match.group(1)), match.group(3), match.group(5))) return out diff --git a/tests/architecture/test_the_register_says_what_holds_a_row.py b/tests/architecture/test_the_register_says_what_holds_a_row.py index 39c52423..e5998482 100644 --- a/tests/architecture/test_the_register_says_what_holds_a_row.py +++ b/tests/architecture/test_the_register_says_what_holds_a_row.py @@ -149,6 +149,43 @@ def test_a_name_the_register_does_not_carry_is_not_invented() -> None: assert register_levels.level_of(_CLI_RUN, "test_absent") == "structural" +def test_the_page_says_what_the_tree_says() -> None: + """The column is a claim, and it decays the way the last table did. + + docs/architecture/typing-scope.md carried a module at 17 errors while + the tree said 12 for weeks: nobody was wrong at the time it was + written. A number in prose is only as fresh as its last reader, so + this reads the page and the tree together and fails when they part. + """ + text = (_ROOT / "docs" / "architecture" / "risk-register.md").read_text(encoding="utf-8") + measured: dict[int, set[str]] = {} + for number, _, level in register_levels.measure(_ROOT, text): + measured.setdefault(number, set()).add(level) + + disagreements = [] + for number, status, _ in register_levels.rows(text): + if status != "closed": + continue + declared = register_levels.declared_level(text, number) + found = "/".join(sorted(measured.get(number, set()))) + if declared != found: + disagreements.append(f"row {number}: page says {declared!r}, tree says {found!r}") + assert not disagreements, ( + f"{disagreements}. Run scripts/register_levels.py and write what it returns." + ) + + +def test_a_row_that_is_not_closed_declares_no_level() -> None: + """Open and partly rows have nothing measured, so a level would be prose.""" + text = (_ROOT / "docs" / "architecture" / "risk-register.md").read_text(encoding="utf-8") + wrong = [ + number + for number, status, _ in register_levels.rows(text) + if status != "closed" and register_levels.declared_level(text, number) != "-" + ] + assert not wrong, f"rows claiming a level with nothing behind them: {wrong}" + + def test_the_measurement_covers_every_closed_reference() -> None: """A classifier that silently skips rows would report a clean page.""" text = (_ROOT / "docs" / "architecture" / "risk-register.md").read_text(encoding="utf-8") From eb582fb85266d5c386911971c91fb633ce081550 Mon Sep 17 00:00:00 2001 From: Evgeny Kiriyak <224408464+evkir@users.noreply.github.com> Date: Tue, 22 Sep 2026 22:39:39 +0300 Subject: [PATCH 3/4] fix(register): no row leaves the table unread, and row 4 names its risk The gap the last commit recorded is closed. A row the pattern cannot parse is absent from the page side and the tree side together, so the comparison could not see it: rows 20 and 24 carry a slash and would have vanished silently under a narrower pattern. The pipe lines are now split into heads, rules and body, and a body line the pattern misses is reported. Counting rows against a fixed number would have been a claim about today instead. Row 4 named test_the_manifest_carries_the_profile, which asserts two keys are present in the manifest. The risk is that two runs seeing different surfaces published the same provenance, and the test beside it -- test_two_profiles_do_not_share_a_hash -- is the one that fails when they do. Both have existed since 2026-09-08; the row named the weaker of the two. The level does not change, which is the point: entrypoint says how the test runs, not whether it asks the right question. That limit is deliberate and now has a name. Restoring the old reference under mutation leaves the suite green, because no measurement here reads what a risk means. The column narrows where to look; it does not judge. --- README.md | 2 +- blog/launch-post-draft.md | 2 +- docs/architecture/risk-register.md | 2 +- scripts/register_levels.py | 28 +++++++++++++++++++ ...test_the_register_says_what_holds_a_row.py | 24 ++++++++++++++++ 5 files changed, 55 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 99f3003b..85d25714 100644 --- a/README.md +++ b/README.md @@ -5,7 +5,7 @@ ![Python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue) ![License](https://img.shields.io/badge/license-Apache_2.0-blue) ![Version](https://img.shields.io/badge/version-v1.7.0-brightgreen) -![Tests](https://img.shields.io/badge/tests-2981%20collected-brightgreen) +![Tests](https://img.shields.io/badge/tests-2983%20collected-brightgreen) ![Mypy](https://img.shields.io/badge/mypy-strict%3A%20104%2F172%20modules-blue) ![LLM](https://img.shields.io/badge/LLM-OpenAI%20%7C%20Anthropic%20%7C%20Ollama-blueviolet) ![Air-Gapped](https://img.shields.io/badge/air--gapped-ready-success) diff --git a/blog/launch-post-draft.md b/blog/launch-post-draft.md index aa7c5e9a..8017294a 100644 --- a/blog/launch-post-draft.md +++ b/blog/launch-post-draft.md @@ -13,7 +13,7 @@ not exist yet, it says so. CyberAI is a multi-agent offensive-security platform: eight agents (recon, intel, exploit, report, planner, mcp-scan, redteam, web3) run a typed, audited pipeline over a shared knowledge base. -2981 tests collected under the gated selection run before every commit, with the +2983 tests collected under the gated selection run before every commit, with the slow and smoke tests deselected there and run separately, `mypy --strict` clean over 104 of 172 modules, Apache-2.0. diff --git a/docs/architecture/risk-register.md b/docs/architecture/risk-register.md index d5fe168e..feb9ac96 100644 --- a/docs/architecture/risk-register.md +++ b/docs/architecture/risk-register.md @@ -49,7 +49,7 @@ whose guard could be stronger, now readable as such without opening the test. | 1 | Web exploitation unreachable under every `--scope` value: two gates normalised opposite sides of one comparison | closed | value | `tests/unit/test_scope_guard_normalises_the_target.py::test_both_gates_reach_the_same_verdict` | | 2 | A skipped phase printed as success, exit code 0 on `state: failed` | closed | entrypoint | `tests/unit/test_a_scope_refusal_is_visible_at_the_end.py::test_the_status_line_reports_partial_for_a_refused_walk` | | 3 | Tests feed argument shapes production never sends | closed | structural | `tests/architecture/test_a_url_parameter_is_fed_a_url.py::test_no_test_feeds_a_url_parameter_something_without_a_scheme` | -| 4 | The published score came from a surface profile blind on both live targets | closed | entrypoint | `tests/unit/test_the_manifest_records_the_surface_it_saw.py::test_the_manifest_carries_the_profile` | +| 4 | The published score came from a surface profile blind on both live targets | closed | entrypoint | `tests/unit/test_the_manifest_records_the_surface_it_saw.py::test_two_profiles_do_not_share_a_hash` -- the row named `test_the_manifest_carries_the_profile` until 2026-09-22, which asserts the keys are present. The risk is that two runs seeing different surfaces published the same provenance, and the test beside it is the one that fails when they do | | 5 | An AI-native benchmark measured a path that called no model | closed | entrypoint | `tests/unit/test_the_bench_table_says_who_solved_it.py::test_a_model_free_run_says_so_next_to_the_score` | | 6 | Documented out-of-band cost understated sixteenfold | closed | value | `tests/unit/test_web_exploit_oob.py::test_every_delivery_is_charged_to_the_budget` | | 7 | One of four web3 engines decided the verdict; aderyn returned nothing | closed | value | `tests/unit/test_web3_agent_merge.py::test_run_merges_and_cross_validates` | diff --git a/scripts/register_levels.py b/scripts/register_levels.py index 94a3bd97..e76ddebc 100644 --- a/scripts/register_levels.py +++ b/scripts/register_levels.py @@ -185,6 +185,34 @@ def rows(text: str) -> list[tuple[int, str, str]]: return out +def table_lines(text: str) -> tuple[list[str], list[str], list[str]]: + """(heads, rules, body) for every line of the page that opens with a pipe. + + A row the pattern cannot parse leaves both sides of the comparison at + once: it carries no declared level and contributes no measured one, so + disagreement is impossible and the row goes unreported. Splitting the + pipe lines three ways is what makes a silent loss visible. + """ + heads, rules, body = [], [], [] + for line in text.splitlines(): + if not line.startswith("|"): + continue + stripped = line.strip() + if stripped.startswith("| # |"): + heads.append(line) + elif set(stripped) <= set("|-"): + rules.append(line) + else: + body.append(line) + return heads, rules, body + + +def unparsed_rows(text: str) -> list[str]: + """Body lines the row pattern does not match.""" + _, _, body = table_lines(text) + return [line for line in body if not _ROW.match(line)] + + def declared_level(text: str, number: int) -> str: """What the page says holds one row, or the empty string if it says nothing.""" for found, _, level in rows(text): diff --git a/tests/architecture/test_the_register_says_what_holds_a_row.py b/tests/architecture/test_the_register_says_what_holds_a_row.py index e5998482..5f182104 100644 --- a/tests/architecture/test_the_register_says_what_holds_a_row.py +++ b/tests/architecture/test_the_register_says_what_holds_a_row.py @@ -186,6 +186,30 @@ def test_a_row_that_is_not_closed_declares_no_level() -> None: assert not wrong, f"rows claiming a level with nothing behind them: {wrong}" +def test_no_row_of_the_table_goes_unread() -> None: + """The gap the previous commit recorded instead of closing. + + A row the pattern misses is absent from the page side and the tree side + together, so the comparison above cannot see it. Counting parsed rows + against a fixed number would be a claim about today; this asks instead + that nothing between the heads and the rules is left over. + """ + text = (_ROOT / "docs" / "architecture" / "risk-register.md").read_text(encoding="utf-8") + heads, rules, body = register_levels.table_lines(text) + assert heads, "no table head -- the page shape changed" + assert len(heads) == len(rules), f"{len(heads)} heads against {len(rules)} rules" + unparsed = register_levels.unparsed_rows(text) + assert not unparsed, f"rows the pattern cannot read: {[line[:60] for line in unparsed]}" + assert len(register_levels.rows(text)) == len(body) + + +def test_a_row_the_pattern_cannot_read_is_reported() -> None: + """The check has to be able to say no, or it says nothing.""" + good = "| # | Risk | Status | Held by | Measured by |\n|---|---|---|---|---|\n" + assert register_levels.unparsed_rows(good + "| 1 | r | closed | value | `x` |\n") == [] + assert register_levels.unparsed_rows(good + "| 1 | r | closed | `x` |\n") + + def test_the_measurement_covers_every_closed_reference() -> None: """A classifier that silently skips rows would report a clean page.""" text = (_ROOT / "docs" / "architecture" / "risk-register.md").read_text(encoding="utf-8") From 220faadd7a55a17bea3c268226f100d45c22d3e6 Mon Sep 17 00:00:00 2001 From: Evgeny Kiriyak <224408464+evkir@users.noreply.github.com> Date: Tue, 22 Sep 2026 22:45:30 +0300 Subject: [PATCH 4/4] docs: the changelog and risk 22 record what the register now checks The release notes carry the column under Added. Risk 22 -- one developer, measuring by hand -- said its mitigation was that every closed row names a test rather than a memory. That is still true and now says more: the row names what kind of test, measured rather than asserted. Neither makes a second pair of eyes out of one, and the row says so; both shorten how long a row can be wrong without anyone noticing. --- CHANGELOG.md | 13 +++++++++++++ docs/architecture/risk-register.md | 2 +- 2 files changed, 14 insertions(+), 1 deletion(-) diff --git a/CHANGELOG.md b/CHANGELOG.md index 161257aa..0ec565fc 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -16,6 +16,19 @@ All notable changes to CyberAI are documented here. icon on a CDN is ordinary and is not flagged; a PNG data URI is inline and is not flagged either. +- **Every closed row of the risk register says what kind of test holds it.** + Risk 20 stood closed for fifteen days on a test that read a config field + while an unscoped run spent fifty-one seconds on a protected range. The + reference resolved the whole time, so the guard over that page was green + and right to be: it asks whether the named test exists. A `Held by` column + now says whether the test asserts about calls, drives a command end to + end, reads a value back, or reads the tree instead of the product -- + written by `scripts/register_levels.py`, compared against the tree by a + test, and reported in CI. No level is forbidden. Two probes established + that the distinction cannot be made from syntax, so `value` marks a row + whose guard could be stronger rather than a defect. Measured on the page + today: structural 2, boundary 2, entrypoint 6, value 6. + - **The capability set a server declares reaches a stage.** The probe had recorded it since it was written and every analysis took tools, transport or a connection flag, so a target's declared surface was collected and diff --git a/docs/architecture/risk-register.md b/docs/architecture/risk-register.md index feb9ac96..f420476d 100644 --- a/docs/architecture/risk-register.md +++ b/docs/architecture/risk-register.md @@ -82,6 +82,6 @@ whose guard could be stronger, now readable as such without opening the test. | 19 | Publishing numbers produced by a path that was broken | closed | structural | `tests/architecture/test_nothing_publishes_without_the_checks.py::test_the_upload_cannot_run_before_the_guard` | | 20 | An offensive tool that does not require an authorisation scope | closed | boundary/value | `tests/unit/test_the_refusal_comes_before_the_first_packet.py::test_an_unscoped_run_never_reaches_the_recon_agent` holds the refusal ahead of the pipeline: measured on 2026-09-17, an unscoped run against a protected range had already spent 51 seconds on nmap, whois, dns and subdomain enumeration before the exploit phase declined it. The row said closed while the target was being touched, because the test behind it read a config field rather than the network. `--no-strict-scope` and `CYBERAI_STRICT_SCOPE=0` remain the named ways to proceed anyway, and since 2026-09-19 they are the only ones: `tests/unit/test_strict_scope.py::test_only_a_named_word_turns_the_refusal_off` holds the refusal against a value nobody chose. The flag reader answers "is this one of the words for yes", so an empty variable -- what a shell leaves for `VAR=` in a .env file -- and a misspelling both read as off, which disarmed the one flag here whose default protects the run | | 21 | Private plans and journals reaching a public repository | unguarded | - | no test and no workflow step checks this. A tree-wide history search finds none of those filenames, which is a measurement of the past, not a guard on the next commit | -| 22 | One developer, measuring by hand | open | - | structural. The mitigation is that every closed row above names a test rather than a memory | +| 22 | One developer, measuring by hand | open | - | structural. The mitigation is that every closed row above names a test rather than a memory, and since 2026-09-22 says what kind of test it is, measured rather than asserted. Neither makes a second pair of eyes out of one; both shorten how long a row can be wrong without anyone noticing | | 23 | A measurement that cannot tell a quiet environment from a clean tree | partly | - | `tests/architecture/test_the_drift_report_refuses_an_unmeasured_environment.py::test_a_checker_that_never_ran_is_not_a_clean_package` and the four rows beside it. Measured on 2026-09-18 on an untouched checkout: with no mypy installed the drift report printed 172 clean modules, 0 errors and no drift and exited zero; without `types-networkx` it printed 275 errors and accused `cyberai/core/kb_graph.py` of being undeclared; on mcp 1.28.1 it printed 284 instead of 285. Five signals, one tree, every guard green, because the guards read declarations rather than the environment. The report now refuses a run it cannot vouch for and names the package that differs. Partly: the same question is unanswered for the test suite, which collected nothing on a machine without `pytest-asyncio` and said so as three collection errors rather than as an unmeasured environment | | 24 | Configuration naming a provider that does not exist | closed | entrypoint/value | `tests/unit/test_the_config_cannot_name_an_unknown_provider.py::test_an_unknown_provider_does_not_reach_the_config` for the environment and `tests/unit/test_the_cli_cannot_name_an_unknown_provider.py::test_an_unknown_provider_is_refused_by_the_parser` for the flag. `CYBERAI_LLM_PROVIDER=gemini` used to produce a config whose provider was that string, and `api_key_for` then resolved a credential by name: a run asked for one vendor could be sent to another under a key nobody named. Both readers narrow against the declared Literal, not a second list -- `test_the_reader_reads_the_declared_literal_and_not_a_copy` fails if a copy appears. The two boundaries answer differently on purpose: the environment falls back to the default because a stale variable must not abort a scan, the flag is refused because whoever typed it is there to read the reply. The three files carrying this path -- and the validator beside them -- entered `[tool.mypy] files` in the same commits, which is what makes the assignments visible to CI at all |