Skip to content

fix: the screen shows what governs the run - #290

Merged
evkir merged 4 commits into
mainfrom
fix/the-screen-shows-what-governs-the-run
Sep 20, 2026
Merged

evkir merged 4 commits into
mainfrom
fix/the-screen-shows-what-governs-the-run

Conversation

@evkir

@evkir evkir commented Sep 20, 2026

Copy link
Copy Markdown
Owner

Four commits, each removing a cause rather than a symptom.

  1. One registry decides which binaries the platform names. The CLI display
    and the probe registry each wrote out the same nine tools and had
    drifted: the display said mas-sentry, the string find_mst hands to
    shutil.which, while the probe registry said mst and wrote that name into
    every run manifest. Toolchain drift was reported against a binary no
    resolver would look for.

  2. status --versions names the version of every tool it found, or the
    reason there is none. Off by default: a probe starts a process per tool,
    and status is the command an operator runs when something is already
    wrong.

  3. The panel shows what governs the run. strict_scope has ended runs before
    the first phase since day 50 and was invisible; so were six other fields
    the orchestrator reads. The rule is scanned out of orchestrator.py and
    fails in both directions.

  4. A typo in CYBERAI_STRICT_SCOPE no longer disarms the refusal. Found by
    printing the new line: an empty variable turned the guard off.

Gate green: ruff, mypy (bare, as CI runs it), drift none, 2957 collected.
Mutation: twelve mutants across the four commits, twelve killed; two
survived a first pass and exposed two defective tests, both fixed.

What this changes

How it was measured

Checklist

  • ruff format --check cyberai/ tests/ and ruff check cyberai/ tests/ pass
  • pytest -W ignore::DeprecationWarning -m "not slow and not smoke" passes
  • New behaviour is covered by a test that fails without the change
  • I have read CLA.md and I hereby sign the CLA

The CLI display and the probe registry each wrote out the same nine tools.
They had drifted: the display said mas-sentry, which is the string find_mst
hands to shutil.which, while the probe registry said mst and wrote that name
into every run manifest. regression_gate reported toolchain drift against a
binary no resolver would ever look for.

_TOOLCHAIN is now derived from TOOL_PROBES, so the two guards that already
stood over the display -- that it names every finder in the package, and that
each label is the binary its finder looks for -- now reach the manifest as
well. Measured: restoring the mst spelling kills
test_every_displayed_name_is_the_binary_its_finder_looks_for; dropping a probe
kills test_the_status_display_names_every_finder_that_exists. Neither guard
could see the probe registry before.

Importing the probe registry into the CLI costs 0.0011 s and one module: the
entry point already imported all nine finders directly.
status named nine binaries and said nothing about which build of them would
run, while environment.py has measured exactly that since day 48. The two
halves existed and never met.

--versions runs each located tool's version flag and prints the number beside
the name, or the reason there is none: searchsploit and mas-sentry carry
flag=None, so they report 'version flag not measured' rather than a guess.
Off by default, on the same split the bench CLI makes -- it imports the probe
only when a manifest is asked for, because a probe starts a process per tool.
status is the command an operator runs when something is already wrong, and
nine hung binaries at the manifest's budget would hold that answer for three
minutes. Measured: a binary that never answers costs exactly the timeout,
20.02 s at VERSION_TIMEOUT, 3.0 s at STATUS_VERSION_TIMEOUT.

Both halves of the display come out of one probe pass. A second walk over the
resolvers can disagree with the first, which is the defect the previous commit
removed between the CLI and the run manifest.

Mutation: four mutants, four killed. Two of them survived the first version of
the tests. The cheap path was asserted with no tool installed, so 'no process
started' held whatever the code did -- the shape of GU, a zero measured on a
population where the signal cannot appear. The one-pass claim was asserted by
counting probe calls, which a second walk over the resolvers never touches.
Both tests now make the binary resolvable first.
The panel printed twelve lines and none of them was strict_scope -- the one
control that ends a run before the first phase since day 50. Six more that
the orchestrator reads were equally invisible: the cost budget, the planner,
replan, model routing, web recon, planned redteam. An operator could read the
whole screen and still not know what the next scan would do.

The rule is scanned, not listed: a test reads orchestrator.py and requires a
line for every config field it consumes. It fails in both directions, so a
control added to the orchestrator without a line fails, and a label left
behind by a control the orchestrator stopped reading fails too. Written this
way on a measurement: four fields on CyberAIConfig -- intel, timeout, verbose,
use_lab_dogfood -- are read by nobody, so a rule saying every field would
have demanded a line for levers that move nothing. use_lab_dogfood has three
tests asserting the value reaches the object and no consumer at all.

Strict scope is the only line that names its source, because it is the only
control here that defaults to on: off has to answer who turned it off, while
off-by-default against six flags that are already off answers nothing. The
source is read as set-or-not rather than inferred from the value, since every
reader in config.py returns the default for unset and for unparseable alike.

Mutation: four mutants, four killed, including one that removes the
orchestrator's read of a field to prove the rule fails from that side too.
Found while printing the new status line: CYBERAI_STRICT_SCOPE= turned the
scope refusal off. So did CYBERAI_STRICT_SCOPE=nope. An empty variable is
what a shell leaves behind for VAR= in a .env file or VAR=$MISSING in a
script, so the guard that risk 20 rests on could come down without anyone
deciding anything.

The cause is not a bug in _env_bool. It answers "is this value one of the
words for yes", and every other value is no -- deliberate, asserted by
test_env_bool_falsy_values with the empty string named in the list. For the
twenty-two flags that default to off that rule costs nothing: an
unrecognised value lands where the default already was. strict_scope is the
one flag here whose default is on, so the same rule handed the refusal to
any string at all.

So the fix is at the point of use, not in the shared reader. _env_guard_bool
recognises both lists -- the words for yes and the words for no -- and
treats anything else as nobody having chosen, leaving the default standing.
That matches what the README table and risk 20 already promised: 0 and
--no-strict-scope are the named ways to proceed without a scope. Now they
are the only ways. It still does not raise: from_env has no raise in it by
design, because garbage in a variable must not abort a scan at startup.

README line 298 is rewritten in the same pass. It said "Refuse the exploit
phase", which stopped being true on day 50 when the refusal moved ahead of
the first phase.

Mutation: four mutants, four killed. The fourth changes _env_bool itself the
way this commit deliberately did not, and it fails test_env_bool_falsy_values
in another file -- the measurement that the old semantics were chosen rather
than overlooked.
@codecov

codecov Bot commented Sep 20, 2026

Copy link
Copy Markdown

Codecov Report

❌ Patch coverage is 97.50000% with 1 line in your changes missing coverage. Please review.

Files with missing lines Patch % Lines
cyberai/core/config.py 92.30% 1 Missing ⚠️

📢 Thoughts on this report? Let us know!

@evkir
evkir merged commit ce47750 into main Sep 20, 2026
10 checks passed
@evkir
evkir deleted the fix/the-screen-shows-what-governs-the-run branch September 20, 2026 07:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant