Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
30 changes: 29 additions & 1 deletion CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -4,8 +4,34 @@ All notable changes to CyberAI are documented here.

## [Unreleased]

### Added

- **The MCP scanner reads the icon field, and scores what a client would
execute.** `icons` arrived with protocol revision 2025-11-25 and is text
shown beside a tool's name before any call. The metadata collector named
seven keys and this was not one of them, so a directive carried in an icon
reached no matcher at all -- not for want of a pattern, but because the
text was never collected. Two categories score the carrier rather than the
reference: an active URL scheme, and SVG by declared type or by name. An
icon on a CDN is ordinary and is not flagged; a PNG data URI is inline and
is not flagged either.

- **The capability set a server declares reaches a stage.** The probe had
recorded it since it was written and every analysis took tools, transport
or a connection flag, so a target's declared surface was collected and
dropped. It is now part of the attestation posture, whose own reason says
MCP has no in-protocol capability attestation.

### Fixed

- **The probe does not offer to open a target's URLs.** URL-mode elicitation
is a client capability, not a server one: `ServerCapabilities` has no field
for it, and a scanner that advertises it has agreed to follow the links of
the endpoint it is scanning. The probe advertises none, which until now was
true by structure alone -- the SDK builds form and URL mode from a single
callback with no separate switch, so adding a form prompt would turn on URL
mode in the same line. A test holds it.

- **The typing measurement refuses an environment it cannot vouch for.** The
drift report read mypy through standard output alone, so a machine without
the checker produced no error lines and the report called the whole package
Expand All @@ -19,7 +45,9 @@ All notable changes to CyberAI are documented here.

- **The published counts name the versions that produced them.** `99/172` and
`285` are facts about a tree read by particular packages, not about the tree
alone. The checker and the SDK that measured them are declared in
alone. Both moved to `100/172` and `284` when `cyberai/agents/mcp_scan/agent.py`
lost its last strict error and the drift step reported it as an undeclared
clean module. The checker and the SDK that measured them are declared in
`[tool.cyberai.measurement]`, named on the typing scope page, and checked
against the bounds that admit them.

Expand Down
4 changes: 2 additions & 2 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -5,8 +5,8 @@
![Python](https://img.shields.io/badge/python-3.11%20%7C%203.12%20%7C%203.13%20%7C%203.14-blue)
![License](https://img.shields.io/badge/license-Apache_2.0-blue)
![Version](https://img.shields.io/badge/version-v1.7.0-brightgreen)
![Tests](https://img.shields.io/badge/tests-2939%20collected-brightgreen)
![Mypy](https://img.shields.io/badge/mypy-strict%3A%2099%2F172%20modules-blue)
![Tests](https://img.shields.io/badge/tests-2946%20collected-brightgreen)
![Mypy](https://img.shields.io/badge/mypy-strict%3A%20100%2F172%20modules-blue)
![LLM](https://img.shields.io/badge/LLM-OpenAI%20%7C%20Anthropic%20%7C%20Ollama-blueviolet)
![Air-Gapped](https://img.shields.io/badge/air--gapped-ready-success)

Expand Down
4 changes: 2 additions & 2 deletions blog/launch-post-draft.md
Original file line number Diff line number Diff line change
Expand Up @@ -13,9 +13,9 @@ not exist yet, it says so.
CyberAI is a multi-agent offensive-security platform: eight agents (recon,
intel, exploit, report, planner, mcp-scan, redteam, web3) run a typed, audited
pipeline over a shared knowledge base.
2939 tests collected under the gated selection run before every commit, with the
2946 tests collected under the gated selection run before every commit, with the
slow and smoke tests deselected there and run separately, `mypy --strict`
clean over 99 of 172 modules, Apache-2.0.
clean over 100 of 172 modules, Apache-2.0.

It is not a wrapper that pipes nmap output into a chat model. Three things make
it a different category of tool.
Expand Down
21 changes: 16 additions & 5 deletions cyberai/agents/mcp_scan/agent.py
Original file line number Diff line number Diff line change
Expand Up @@ -26,7 +26,7 @@
from cyberai.core.base_agent import BaseAgent, Tool
from cyberai.core.scan_session import Severity
from cyberai.mcp.auth_metadata import probe_auth_metadata
from cyberai.mcp.client_probe import probe
from cyberai.mcp.client_probe import MCPProbeResult, probe


def _run_coro(coro: Any) -> Any:
Expand Down Expand Up @@ -70,7 +70,9 @@ def _register_tools(self) -> None:

def _probe(self, endpoint: str, transport: Optional[str] = None) -> dict[str, Any]:
"""Run the async probe to completion and return a plain dict."""
result = _run_coro(probe(endpoint, transport)) # type: ignore[arg-type]
# _run_coro is deliberately untyped -- it takes any coroutine -- so the
# probe's own result type is what says this is a dict, not the runner.
result: MCPProbeResult = _run_coro(probe(endpoint, transport)) # type: ignore[arg-type]
return result.to_dict()

def run(self, target: str, context: Optional[dict[str, Any]] = None) -> dict[str, Any]:
Expand Down Expand Up @@ -98,7 +100,11 @@ def run(self, target: str, context: Optional[dict[str, Any]] = None) -> dict[str
overprivilege = self._analyze_overprivilege(target, probe_result["tools"])
exposure = self._assess_exposure(target, probe_result["transport"], probe_result["tools"])
attestation = self._assess_attestation(
target, probe_result["transport"], probe_result["connected"], probe_result["error"]
target,
probe_result["transport"],
probe_result["connected"],
probe_result["error"],
probe_result["capabilities"],
)
trust = self._analyze_trust(target, probe_result["tools"])
auth_metadata = probe_auth_metadata(target, probe_result["transport"]).to_dict()
Expand Down Expand Up @@ -246,15 +252,20 @@ def _assess_exposure(
return {"exposed": scan.is_exposed, "scan": scan.to_dict()}

def _assess_attestation(
self, target: str, transport: str, connected: bool, error: str | None
self,
target: str,
transport: str,
connected: bool,
error: str | None,
capabilities: dict[str, Any] | None = None,
) -> dict[str, Any]:
"""Assess the transport-authentication posture of the target endpoint.

Like exposure this is an endpoint property, so at most one Finding is
recorded, and only when the endpoint accepted an unauthenticated
session. stdio and undetermined remote endpoints produce no Finding.
"""
scan = assess_attestation(target, transport, connected, error)
scan = assess_attestation(target, transport, connected, error, capabilities)
if scan.is_finding:
self.session.add_finding(
severity=Severity(scan.severity),
Expand Down
9 changes: 9 additions & 0 deletions cyberai/agents/mcp_scan/attestation.py
Original file line number Diff line number Diff line change
Expand Up @@ -47,6 +47,7 @@ class AttestationScan:
connected: bool = False
unauthenticated: bool = False
transport_encrypted: bool = True
declared_capabilities: list[str] = field(default_factory=list)
severity: str = Severity.INFO.value
reasons: list[str] = field(default_factory=list)

Expand Down Expand Up @@ -75,6 +76,7 @@ def assess_attestation(
transport: str,
connected: bool,
error: str | None = None,
capabilities: dict[str, Any] | None = None,
) -> AttestationScan:
"""Assess the transport-authentication posture of an MCP endpoint.

Expand All @@ -83,6 +85,13 @@ def assess_attestation(
accepted an anonymous session.
"""
scan = AttestationScan(endpoint=endpoint, transport=transport)
# What the server said it can do, which is the thing the reason below
# calls unattested. The probe has collected this since it was written and
# no stage read it, so the set a server advertises -- experimental blocks,
# protocol extensions, task support -- reached no inventory and no report.
# Names only: the values are free-form per the spec and belong in the raw
# probe dump, not in a posture summary.
scan.declared_capabilities = sorted(capabilities or {})

if transport == "stdio":
scan.reasons.append(
Expand Down
24 changes: 23 additions & 1 deletion cyberai/agents/mcp_scan/poisoning.py
Original file line number Diff line number Diff line change
Expand Up @@ -38,6 +38,21 @@
),
(r"include .{0,30}(api[_ ]?key|token|secret|password|credential)", "credential_harvest"),
(r"read .{0,30}(\.env|id_rsa|/etc/passwd|ssh key|config file)", "sensitive_read"),
# Icons, revision 2025-11-25. What is scored is the carrier, not the fact
# of an external reference: a tool icon served from a CDN is ordinary, and
# this matcher sees flattened text, so it cannot tell the server's own
# origin from anyone else's. A pattern on "src is remote" would therefore
# flag the normal case and measure nothing. These two describe content the
# client executes or renders as markup while showing the tool's name --
# before the user has agreed to call anything.
(
r'"src"\s*:\s*"\s*(?:javascript:|vbscript:|data:text/html|data:image/svg)',
"icon_active_scheme",
),
(
r'"mimeType"\s*:\s*"image/svg\+xml"|"src"\s*:\s*"[^"]+\.svg(?:[?#][^"]*)?"',
"icon_executable_carrier",
),
]
_MCP_COMPILED = [
(re.compile(pat, re.IGNORECASE | re.DOTALL), label) for pat, label in MCP_POISONING_PATTERNS
Expand Down Expand Up @@ -94,7 +109,14 @@ def _collect_text(tool: dict[str, Any]) -> tuple[str, list[str]]:
if schema_strings:
parts.extend(schema_strings)
fields.append("inputSchema")
for key in ("annotations", "meta", "outputSchema"):
# `icons` arrived with revision 2025-11-25 and is server-controlled text
# that reaches the client before any call: `src` is a URL the client is
# expected to fetch, and `mimeType`/`sizes` are free strings rendered next
# to the tool's name. The whitelist above named seven keys and this was not
# one of them, so a directive carried in an icon field reached no matcher
# at all -- not because no pattern described it, but because the text was
# never collected. A channel nothing reads cannot be scored.
for key in ("annotations", "meta", "outputSchema", "icons"):
val = tool.get(key)
if val:
parts.append(json.dumps(val, default=str))
Expand Down
4 changes: 2 additions & 2 deletions docs/architecture/risk-register.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Risk Register — what could make this project fail

**Last verified against the tree:** 2026-09-17.
**Last verified against the tree:** 2026-09-19.

This page replaces a document that tracked eight build-out defects and had
said "ALL RESOLVED" since day 7 while the README went on describing it as the
Expand Down Expand Up @@ -38,7 +38,7 @@ not fixed. `unguarded` — held by discipline, with no machine behind it.
|---|---|---|---|
| 8 | The CVE-Bench adapter was written against criteria older than v2.1.0 | partly | `tests/integration/test_the_bench_answers_to_the_upstream.py::test_the_adapter_answers_to_the_checkout_on_disk` — carries the smoke marker, so CI skips it and only a workstation with the checkout runs it |
| 9 | The MCP client did not report which protocol revision it negotiated | closed | `tests/integration/test_mcp_revision_in_output.py::test_the_terminal_names_the_revision_the_probe_negotiated` |
| 10 | The detector does not cover tool icons or URL-mode elicitation, both added to the protocol in revision 2025-11-25 | open | no test mentions either; a tree-wide search for both terms returns nothing |
| 10 | The detector does not cover tool icons or URL-mode elicitation, both added to the protocol in revision 2025-11-25 | partly | Icons: `tests/unit/test_mcp_poisoning.py::test_a_directive_in_an_icon_field_reaches_the_matcher` and `tests/unit/test_mcp_poisoning.py::test_an_executable_icon_carrier_is_a_signal`. The field was outside the collected whitelist, so no pattern could have reached it. Elicitation: the risk was written on a premise the SDK does not support -- `ServerCapabilities` has no elicitation field, so a scanned server cannot advertise URL mode and a probe that calls nothing never receives error -32042. The exposure runs the other way and is held by `tests/unit/test_the_probe_does_not_offer_to_open_a_url.py::test_the_probe_advertises_no_elicitation`. Partly: what a target does with elicitation is reachable only by calling its tools, which this scanner does not do |
| 11 | Detection and false-positive figures come from a corpus this project wrote | open | `tests/architecture/test_corpus_integrity.py::test_both_classes_meet_the_floor` guards the corpus, not its provenance. No external corpus has ever been run |
| 12 | CVE-Bench and phantom-grid both claim port 9090 | closed | `tests/unit/test_cve_bench_driver.py::test_a_taken_port_is_named_not_blamed_on_the_stack` |
| 13 | Two finding counts in one result, neither reconciled with the other | closed | `tests/unit/test_web3_agent_merge.py::test_aderyn_only_critical_raises_the_headline` |
Expand Down
6 changes: 3 additions & 3 deletions docs/architecture/typing-scope.md
Original file line number Diff line number Diff line change
@@ -1,6 +1,6 @@
# Typing scope

`mypy --strict` reads 99 of 172 modules in the package. The other 73 hold 285
`mypy --strict` reads 100 of 172 modules in the package. The other 72 hold 284
errors and are not checked.

The scope is a list of named modules, so a module that passes strictly stays
Expand All @@ -13,7 +13,7 @@ module that could be declared and is not becomes a failing CI step rather than
a quiet omission.

Not checked is stronger than it sounds, and the boundary is the reason. Of
the 99 modules in the scope, 22 import a module outside it at module level,
the 100 modules in the scope, 22 import a module outside it at module level,
and between them they reach 29 such modules. mypy follows those imports to
resolve names and does not report what it finds there: measured by appending
an unannotated function to `cyberai/core/config.py`, which is outside the
Expand Down Expand Up @@ -121,7 +121,7 @@ the tests installs no stubs at all.

## The unchecked side

Six modules carry roughly a third of the 285 errors:
Six modules carry roughly a third of the 284 errors:

| Module | Errors |
|---|---|
Expand Down
29 changes: 27 additions & 2 deletions docs/redteam/mcp-scanning.md
Original file line number Diff line number Diff line change
Expand Up @@ -31,17 +31,42 @@ landscape.

| Stage | What it looks for | OWASP MCP Top 10 | MITRE ATLAS |
| --- | --- | --- | --- |
| tool-poisoning | Hidden instructions, unicode tricks, base64, hidden HTML in tool metadata | MCP03:2025 Tool Poisoning | AML.T0110 AI Agent Tool Poisoning |
| tool-poisoning | Hidden instructions, unicode tricks, base64, hidden HTML, executable icon carriers in tool metadata | MCP03:2025 Tool Poisoning | AML.T0110 AI Agent Tool Poisoning |
| over-privilege | Tools that touch fs/net/exec beyond their declared purpose | MCP02:2025 Privilege Escalation via Scope Creep | AML.T0086 Exfiltration via AI Agent Tool Invocation |
| trust-propagation | Steering / shadowing of sibling tools, cross-server name collisions | MCP06:2025 Intent Flow Subversion | AML.T0051 LLM Prompt Injection |
| attestation | Anonymous acceptance, self-asserted identity, no message auth | MCP07:2025 Insufficient Authentication & Authorization | - |
| attestation | Anonymous acceptance, self-asserted identity, no message auth, the capability set the server declares | MCP07:2025 Insufficient Authentication & Authorization | - |
| exposure | Remote reachability, DNS-rebinding surface, dangerous capabilities | MCP07:2025 Insufficient Authentication & Authorization | AML.T0040 AI Model Inference API Access |
| mst-fuzzing | Low-level malformed / protocol fuzzing (optional, see below) | MCP05:2025 Command Injection & Execution | AML.T0110 AI Agent Tool Poisoning |

MCP06 is titled *Intent Flow Subversion* in the OWASP index and *Prompt
Injection via Contextual Payloads* in the project README; the taxonomy is in
beta and both names refer to the same category.

### Icons, and the two directions of URL-mode elicitation

Revision 2025-11-25 added `icons` to tools, prompts, resources and the server's
own identity. It is text a client shows beside a tool's name before any call,
which makes it the same channel as `description` and it is scanned as one.

Two categories score it, and both are about the carrier rather than the
reference. An icon served from a CDN is how the field is meant to be used, and
the scanner reads flattened metadata, so it cannot tell the server's own origin
from anyone else's: a rule on "the source is remote" would flag the ordinary
case. What is scored is content a client executes or renders as markup --
`javascript:`, `vbscript:`, `data:text/html`, `data:image/svg`, a declared
`image/svg+xml`, or an `.svg` name. A PNG data URI is inline and is not
flagged; inline is not the property, executable is.

URL-mode elicitation is not a property of the scanned server at all, and the
scanner has no category for it. `ServerCapabilities` has no elicitation field:
the client declares willingness to open a URL, and a server asks for one with
error `-32042` in reply to a tool call. A probe that inventories a surface
calls nothing, so there is nothing to observe. The exposure runs the other
way -- a scanner that advertises the capability has agreed to follow the links
of the endpoint it is scanning -- and the probe therefore advertises no
elicitation at all. The SDK builds form and URL mode from a single callback
with no separate switch, so that is held by a test rather than by care.

## Usage

Inventory a target (transport is inferred from the endpoint):
Expand Down
1 change: 1 addition & 0 deletions pyproject.toml
Original file line number Diff line number Diff line change
Expand Up @@ -107,6 +107,7 @@ files = [
"cyberai/agents/intel/__init__.py",
"cyberai/agents/intel/epss_client.py",
"cyberai/agents/mcp_scan/__init__.py",
"cyberai/agents/mcp_scan/agent.py",
"cyberai/agents/mcp_scan/attestation.py",
"cyberai/agents/mcp_scan/exposure.py",
"cyberai/agents/mcp_scan/mst_bridge.py",
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -89,7 +89,7 @@ def _crossings() -> tuple[set[pathlib.Path], set[pathlib.Path]]:

def test_the_scope_covers_the_modules_it_declares() -> None:
"""The premise the rest of this file argues about."""
assert len(_scope()) == 99
assert len(_scope()) == 100
assert len(list(_PACKAGE.rglob("*.py"))) == 172


Expand Down
Loading
Loading