Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
2 changes: 2 additions & 0 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -164,6 +164,8 @@ jobs:
run: python scripts/typing_scope_drift.py
- name: Check the declared stubs are the stubs that ship
run: python scripts/stub_distributions.py
- name: Report what holds each closed row of the risk register
run: python scripts/register_levels.py

mcp-1x:
name: MCP surface on mcp 1.x
Expand Down
13 changes: 13 additions & 0 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -16,6 +16,19 @@ All notable changes to CyberAI are documented here.
icon on a CDN is ordinary and is not flagged; a PNG data URI is inline and
is not flagged either.

- **Every closed row of the risk register says what kind of test holds it.**
Risk 20 stood closed for fifteen days on a test that read a config field
while an unscoped run spent fifty-one seconds on a protected range. The
reference resolved the whole time, so the guard over that page was green
and right to be: it asks whether the named test exists. A `Held by` column
now says whether the test asserts about calls, drives a command end to
end, reads a value back, or reads the tree instead of the product --
written by `scripts/register_levels.py`, compared against the tree by a
test, and reported in CI. No level is forbidden. Two probes established
that the distinction cannot be made from syntax, so `value` marks a row
whose guard could be stronger rather than a defect. Measured on the page
today: structural 2, boundary 2, entrypoint 6, value 6.

- **The capability set a server declares reaches a stage.** The probe had
recorded it since it was written and every analysis took tools, transport
or a connection flag, so a target's declared surface was collected and
Expand Down
2 changes: 1 addition & 1 deletion README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,7 +5,7 @@
![Python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue)
![License](https://img.shields.io/badge/license-Apache_2.0-blue)
![Version](https://img.shields.io/badge/version-v1.7.0-brightgreen)
![Tests](https://img.shields.io/badge/tests-2970%20collected-brightgreen)
![Tests](https://img.shields.io/badge/tests-2983%20collected-brightgreen)
![Mypy](https://img.shields.io/badge/mypy-strict%3A%20104%2F172%20modules-blue)
![LLM](https://img.shields.io/badge/LLM-OpenAI%20%7C%20Anthropic%20%7C%20Ollama-blueviolet)
![Air-Gapped](https://img.shields.io/badge/air--gapped-ready-success)
Expand Down
2 changes: 1 addition & 1 deletion blog/launch-post-draft.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,7 +13,7 @@ not exist yet, it says so.
CyberAI is a multi-agent offensive-security platform: eight agents (recon,
intel, exploit, report, planner, mcp-scan, redteam, web3) run a typed, audited
pipeline over a shared knowledge base.
2970 tests collected under the gated selection run before every commit, with the
2983 tests collected under the gated selection run before every commit, with the
slow and smoke tests deselected there and run separately, `mypy --strict`
clean over 104 of 172 modules, Apache-2.0.

Expand Down
86 changes: 54 additions & 32 deletions docs/architecture/risk-register.md

Large diffs are not rendered by default.

249 changes: 249 additions & 0 deletions scripts/register_levels.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,249 @@
"""What kind of assertion holds each closed row of the risk register.

Risk 20 stood closed for fifteen days on a test that read a config field
while an unscoped run spent fifty-one seconds touching a protected range.
The reference resolved, so the existing guard was green and right to be:
it asks whether the named test exists, not what the test asserts.

Two probes on 2026-09-22 established that "measures behaviour" cannot be
decided syntactically. A test calling validate_exploit_scope and checking
its verdict is indistinguishable from one calling from_env and reading
.strict_scope; a rule that fails the second fails the first as well, and
that accusation is false. So this reports rather than judges. The register
declares what holds each row, this measures it, and the guard beside it
fails when the two disagree. A row held by a value read is not forbidden.
It is visible, which is the whole of the fix.

Levels, in the order they are decided:

structural the test imports nothing from cyberai; it reads the tree,
a workflow file or a document. Deliberate for some rows.
boundary something asserts about calls: assert_not_called, call_count,
call_args. The strongest form available here.
entrypoint a run-shaped call carries something from cyberai as an
argument: CliRunner().invoke(cli, ...), asyncio.run(probe()).
The receiver is the runner, never the product, so the
argument is what says the product was driven end to end.
Measured 2026-09-22: this catches the six CLI rows and
leaves the config reads alone. It does NOT yet catch a run
on an object a helper built -- rows 7, 13 and 20 read as
value for that reason. A level below the truth is not a
false claim, and the guard beside this compares what the
register declares against what this returns.
value the product is exercised and the assertions are about values
it returned or holds.

Helper functions in the same module are followed, because a test whose
body is three calls to module helpers says nothing about itself.
"""

from __future__ import annotations

import ast
import collections
import pathlib
import re
import sys

_ROOT = pathlib.Path(__file__).resolve().parents[1]
_REGISTER = _ROOT / "docs" / "architecture" / "risk-register.md"

_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|\s*([\w/-]+)\s*\|(.+)\|\s*$")
_REFERENCE = re.compile(r"tests/[A-Za-z0-9_/]+\.py::[A-Za-z0-9_]+")

_CALL_ASSERTION = re.compile(r"assert_(?:not_)?(?:called|awaited)\w*|call_count|call_args")
_RUN_NAMES = frozenset({"invoke", "run"})

LEVELS = ("structural", "boundary", "entrypoint", "value")


def product_names(tree: ast.Module) -> set[str]:
"""Names this module pulled out of cyberai, under whatever alias."""
names: set[str] = set()
for node in ast.walk(tree):
if isinstance(node, ast.ImportFrom):
if (node.module or "").startswith("cyberai"):
names |= {alias.asname or alias.name for alias in node.names}
elif isinstance(node, ast.Import):
for alias in node.names:
if alias.name.startswith("cyberai"):
names.add(alias.asname or alias.name.split(".")[0])
return names


def _root_name(node: ast.expr) -> str | None:
while isinstance(node, (ast.Attribute, ast.Subscript, ast.Call)):
node = node.value if isinstance(node, (ast.Attribute, ast.Subscript)) else node.func
return node.id if isinstance(node, ast.Name) else None


def _argument_names(call: ast.Call) -> set[str]:
"""Every name and attribute appearing in the arguments of one call."""
out: set[str] = set()
for argument in [*call.args, *(keyword.value for keyword in call.keywords)]:
for child in ast.walk(argument):
if isinstance(child, ast.Name):
out.add(child.id)
elif isinstance(child, ast.Attribute):
out.add(child.attr)
return out


def _mentioned(node: ast.AST) -> set[str]:
out: set[str] = set()
for child in ast.walk(node):
if isinstance(child, ast.Name):
out.add(child.id)
elif isinstance(child, ast.Attribute):
out.add(child.attr)
return out


def _functions(tree: ast.Module) -> dict[str, ast.FunctionDef | ast.AsyncFunctionDef]:
return {
node.name: node
for node in ast.walk(tree)
if isinstance(node, (ast.FunctionDef, ast.AsyncFunctionDef))
}


def level_of(source: str, function: str) -> str:
"""The level of one test, following helpers defined in the same module."""
tree = ast.parse(source)
product = product_names(tree)
functions = _functions(tree)
if function not in functions:
return "structural"

seen: set[str] = set()
queue: collections.deque[str] = collections.deque([function])
reaches = False
boundary = False
entrypoint = False

while queue:
current = queue.popleft()
if current in seen:
continue
seen.add(current)
node = functions.get(current)
if node is None:
continue
segment = ast.get_source_segment(source, node) or ""
if _CALL_ASSERTION.search(segment):
boundary = True
for call in ast.walk(node):
if not isinstance(call, ast.Call):
continue
if isinstance(call.func, ast.Name):
called = call.func.id
elif isinstance(call.func, ast.Attribute):
called = call.func.attr
root = _root_name(call.func.value)
if root is not None and root in product:
reaches = True
else:
continue
if called in product:
reaches = True
if called in _RUN_NAMES and _argument_names(call) & product:
entrypoint = True
reaches = True
if product & _mentioned(node):
reaches = True
for name in _mentioned(node):
if name in functions and name not in seen:
queue.append(name)

if not reaches:
return "structural"
if boundary:
return "boundary"
if entrypoint:
return "entrypoint"
return "value"


def closed_references(text: str) -> list[tuple[int, str]]:
"""(row number, node id) for every reference a closed row carries."""
out: list[tuple[int, str]] = []
for line in text.splitlines():
match = _ROW.match(line)
if match and match.group(3) == "closed":
for reference in _REFERENCE.findall(match.group(5)):
out.append((int(match.group(1)), reference))
return out


def rows(text: str) -> list[tuple[int, str, str]]:
"""(number, status, declared level) for each numbered row."""
out: list[tuple[int, str, str]] = []
for line in text.splitlines():
match = _ROW.match(line)
if match:
out.append((int(match.group(1)), match.group(3), match.group(4)))
return out


def table_lines(text: str) -> tuple[list[str], list[str], list[str]]:
"""(heads, rules, body) for every line of the page that opens with a pipe.

A row the pattern cannot parse leaves both sides of the comparison at
once: it carries no declared level and contributes no measured one, so
disagreement is impossible and the row goes unreported. Splitting the
pipe lines three ways is what makes a silent loss visible.
"""
heads, rules, body = [], [], []
for line in text.splitlines():
if not line.startswith("|"):
continue
stripped = line.strip()
if stripped.startswith("| # |"):
heads.append(line)
elif set(stripped) <= set("|-"):
rules.append(line)
else:
body.append(line)
return heads, rules, body


def unparsed_rows(text: str) -> list[str]:
"""Body lines the row pattern does not match."""
_, _, body = table_lines(text)
return [line for line in body if not _ROW.match(line)]


def declared_level(text: str, number: int) -> str:
"""What the page says holds one row, or the empty string if it says nothing."""
for found, _, level in rows(text):
if found == number:
return level
return ""


def measure(root: pathlib.Path, text: str) -> list[tuple[int, str, str]]:
"""(row, reference, level) for every closed reference in the register."""
out: list[tuple[int, str, str]] = []
for number, reference in closed_references(text):
relative, _, function = reference.partition("::")
path = root / relative
if not path.exists():
out.append((number, reference, "structural"))
continue
out.append((number, reference, level_of(path.read_text(encoding="utf-8"), function)))
return out


def main() -> int:
measured = measure(_ROOT, _REGISTER.read_text(encoding="utf-8"))
counts: collections.Counter[str] = collections.Counter(level for _, _, level in measured)
for number, reference, level in measured:
print(f"{number:>3} {level:<11} {reference}")
print()
print(" ".join(f"{level}: {counts[level]}" for level in LEVELS))
print(f"references: {len(measured)}")
return 0


if __name__ == "__main__":
sys.exit(main())
12 changes: 9 additions & 3 deletions tests/architecture/test_the_register_names_tests_that_exist.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,17 +26,23 @@
_REGISTER = _ROOT / "docs" / "architecture" / "risk-register.md"

_REFERENCE = re.compile(r"tests/[A-Za-z0-9_/]+\.py::[A-Za-z0-9_]+")
_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|(.+?)\|\s*$")
_ROW = re.compile(r"^\|\s*(\d+)\s*\|(.+?)\|\s*(\w+)\s*\|\s*([\w/-]+)\s*\|(.+?)\|\s*$")
_DECLARED = {"closed", "partly", "open", "unguarded"}


def rows(text: str) -> list[tuple[int, str, str]]:
"""(number, status, evidence cell) for each numbered row in the register."""
"""(number, status, evidence cell) for each numbered row in the register.

The table grew a fourth column on 2026-09-22 saying what kind of
assertion holds each closed row. Both readers of this page matched it
into the evidence cell and kept working, which is the wrong reason for
a check to be green: the pattern has to know the shape it reads.
"""
out = []
for line in text.splitlines():
match = _ROW.match(line)
if match:
out.append((int(match.group(1)), match.group(3), match.group(4)))
out.append((int(match.group(1)), match.group(3), match.group(5)))
return out


Expand Down
Loading
Loading