Skip to content

Repository files navigation

NetSecOps

Read-only configuration & vulnerability assessment for network and security infrastructure

Python 3.12+ FastAPI React 18 PostgreSQL 16 Read-only Phases 0-6 and 8 complete, 7 in progress MIT


Contents

What NetSecOps is The one-paragraph version, and the guarantee it is built around
Architecture at a glance Five diagrams: the components, the read-only guard, the pipeline, path analysis, deployment
Getting started Running it locally in about five minutes, with something to look at
Where the project stands Which phases are finished, and what the unfinished one still owes
What it does, phase by phase The capabilities in detail, oldest to newest
What building it taught us Three defects that produced confident answers rather than errors
Repository layout Where things live, and the files worth reading first
Development Tests, linting, conventions
The command line netsecops-cli, including break-glass recovery
Smoke-testing a deployment Checking a running install end to end
Documentation The SRS, deployment, evaluation and commercial docs

What NetSecOps is

NetSecOps is a self-hosted, web-based platform that assesses the configuration and vulnerability posture of network and security devices — Cisco, Palo Alto Networks, Fortinet and Check Point — and tracks how that posture changes over time.

It authenticates to devices over SSH and vendor HTTPS APIs, extracts running configuration and operational state, normalises it into a vendor-neutral model, and evaluates it against a library of hardening, firewall-hygiene, AAA and crypto checks, correlating software versions against CVE and vendor PSIRT advisories.

It never changes a target device

This is a hard constraint, not a policy setting (SRS §8):

  • Allow-list, not deny-list. Every vendor adapter declares the exact set of commands and API operations it may issue. Anything else is rejected before transmission.
  • Defence in depth. A global deny-list additionally blocks write verbs (configure, write, copy, reload, set, commit, …) even if an allow-list entry were mis-specified.
  • GET-only REST. POST is permitted solely for authentication and for vendor APIs that are POST-only by design (Check Point show-*, FortiManager method: "get").
  • No side effects. Adapters never run ping, test aaa, debug, or anything that writes a file or sends a packet from the device.
  • No unchecked path. Adapters hold a guarded session, not a transport. An adapter author cannot forget to check, because there is nothing unchecked to reach for.
  • Proven in CI. 283 conformance assertions check what the guard decides, and a fake SSH device that records every byte it receives checks what actually arrives. The build fails on either.
  • Transparent. Every command sent to a device is recorded in a tamper-evident audit log, so customers can see exactly what ran.

Architecture at a glance

Five pictures, in the order they answer the questions people actually ask. All of them are generated by one script with no dependency beyond Pillow — see docs/diagrams/ to regenerate them.

What the pieces are

Logical architecture

A React console behind Caddy, a FastAPI control plane, PostgreSQL, and a pool of workers. The services layer holds the business logic and knows nothing about HTTP, which is why the same code serves a request and runs a scheduled job. Workers are the only thing in the system that opens a connection to a device.

How the read-only guarantee is enforced

The read-only guarantee

The four layers described above, drawn in the order a command meets them. Every one of them acts before transmission — nothing is filtered after the fact, and there is no path around them, because an adapter holds a guarded session rather than a transport.

How a device becomes a finding

From a device to a finding

Collection, a sealed artefact, one parser per platform, and then the normalised configuration model that every engine reads. Nothing downstream of that model parses vendor syntax, which is what lets a check be written once and evaluated against thirteen platforms.

Why a path answer has two verdicts

A path answer has two axes

Routing and policy fail for independent reasons, so NetSecOps reports them separately rather than collapsing them into one word. The example is the demonstration estate, and it is the real answer — a test pins it: the packet is routed end to end and every firewall permits it, but one of them translates addresses in a way the trace cannot follow, so the firewalls after it were asked about the wrong ones. NAT that can be followed is followed, and does not weaken the verdict; this estate's edge rule translates to whatever address an interface happens to hold, which the rule never states. Somebody opens a firewall on this answer, so it says what it does not know.

What you actually run

What actually runs

One Docker Compose stack on your own infrastructure, with outbound connections only to the devices you name, only on tcp/22 and tcp/443, and only read-only. A hundred-device install and a two-thousand-device install run the same product.


Getting started

Requirements: Docker 24+ and Docker Compose. Nothing else.

git clone https://github.com/Krishcalin/NetSecOps.git
cd NetSecOps

make up              # generates keys, builds, migrates, prints the URL
make create-admin    # create the first Super Admin
make demo-seed       # optional: a demonstration estate, so there is something to look at

Then open http://localhost:8080.

make demo-seed stands up four devices across three vendors, ingests a configuration for each and assesses them — no device is contacted, no credential is needed, and the findings are the product's real opinion of those configurations rather than fixtures. It refuses to run if the inventory already holds a device it did not create, and make demo-purge removes exactly what it created. What to look at, and in what order.

Back up MASTER_KEY separately from the database. It wraps every stored device credential. Losing it means losing them all; storing it beside a database dump means a single stolen backup yields both.

Port already in use? Every published port is overridable in .env, which matters on a workstation running several projects:

UI_PORT=8088     # SPA           (default 8080)
API_PORT=8010    # API           (default 8000)
DB_PORT=5442     # PostgreSQL    (default 5442, chosen to avoid a local 5432)

Running it without Docker

make install         # backend venv + frontend node_modules
make db              # just PostgreSQL, on host port 5442
make migrate
make dev-api         # http://localhost:8000
make dev-ui          # http://localhost:5173

It can be evaluated without a device

A product nobody can reach and a product nobody can try are the same failure at different scales. This category's normal answer to "can I see it" is a nine-to-thirteen week implementation — credentials brokered, firewall rules opened, collectors sited, a change window found. Nobody evaluates a tool on that budget.

netsecops-cli demo-seed stands up four devices across three vendors, ingests a configuration for each, assesses them, imports advisories and matches them. It takes about ten seconds and contacts nothing.

The devices are not real; everything said about them is. The configurations are ingested through the same path an operator's upload uses (FR-COL-11), parsed by the same parsers, sealed as the same artefacts, and assessed by the same engine against the same 104-check library. There is no demonstration write path — a demonstration write path is how a demo comes to show something the product does not do.

The estate is designed rather than sampled, so that each thing this product does differently has something real to show: an ordinary neglected switch for the findings, a router with no rulebase that reports no decision rather than "allowed", a firewall that permits and translates so the path verdict is partially-allowed with the translating device named, a rulebase carrying a shadowed rule and an any-any permit, and a software version old enough for the vulnerability engine to match.

It refuses to run if the inventory holds a device it did not create, every device it makes carries the tag netsecops-demo, and demo-purge removes exactly those — leaving alone any device added by hand, because that command runs at the moment somebody is onboarding their first real one.

docs/evaluating.md says what to look at and in what order, including the limits the demonstration will show you.

Building it found four defects in path analysis that no test had caught, all of them false-blocked verdicts — the dangerous direction, because a blocked verdict says a control is already in place and somebody stops looking. One of them left every ASA rulebase unable to match anything at all. They are described under Phase 8 and in What building it taught us.

Where the project stands

Development follows the phase plan in SRS §12. Phases 0–5 were built strictly in order, each one's acceptance criteria passing before the next began.

Phases 6, 7 and 8 were opened at the same time, which is a deliberate departure from that rule and is recorded in SRS §12 rather than left implicit. Phases 6 and 8 have since met their acceptance criteria. Phase 7 has not, and its section says exactly which part is missing — a phase that is 80% done is far easier to misread as finished than one that has not started.

Phase 8 was not in the SRS as issued. It was added after a competitive analysis found that multi-device reasoning — "can this host reach that one, and what decides" — is the one capability separating this product from the established tools in its category, and that most of the other gaps identified collapse into it.

Phase Scope Status
0 Monorepo, auth/MFA/RBAC, credential vault, audit chain, CI, Docker Complete
1 Inventory, credentials, job engine, read-only enforcement framework Complete
2 Cisco IOS/IOS-XE/NX-OS/ASA collection, parsing, drift Complete
3 Check engine + baseline library, findings, compliance mapping Complete
4 Palo Alto, Fortinet, Check Point + firewall rulebase analysis Complete
5 Wireless (WLC/9800) + AAA: ISE, FortiAuthenticator, FreeRADIUS, tac_plus Complete
6 Vulnerability assessment: NVD, CSAF, PSIRT, EoL, KEV/EPSS Complete — acceptance met
7 Discovery, reporting, integrations, hardening In progress — every M requirement built; TEST-08 needs a physical device lab
8 Topology and path analysis Complete — acceptance met

Thirteen platforms are collected and parsed, and the check library stands at 104.


What it does, phase by phase

Each phase below says what it delivers and, where relevant, what it still owes. They are in order; the newest work is Phases 6 to 8 at the end.

Phase 0 — foundations

  • Authentication — Argon2id password hashing, a configurable password policy with a no-reuse history window, and account lockout after repeated failures.
  • Sessions — short-lived access tokens (≤15 min) and rotating, revocable refresh tokens (≤8 h), delivered as Secure; HttpOnly; SameSite=Strict cookies. Replaying a rotated refresh token revokes the whole session family and raises an audit event.
  • MFA — RFC 6238 TOTP with single-use recovery codes; codes cannot be replayed inside their validity window.
  • Single sign-on (FR-AUTH-04) — OIDC Authorization Code with PKCE, and three positions worth stating because each had a more convenient alternative. The provider says who is signing in, not who may: an assertion for a subject with no account is refused and audited, never provisioned, because a directory group is a statement about employment and not an authorisation grant. Roles follow an administrator's mapping, which the console edits — group membership drives the roles the mapping names, so leaving a group takes the role away, while a role granted by hand and named in no mapping survives every login; Super Admin cannot be mapped at all, being the role that can rewrite the mapping. It replaces the password, not the second factor: an enrolled user is still asked for their TOTP, because the provider may have performed its own MFA or may be checking one directory password, and the assertion does not reliably say which. The sign-in state lives in a table rather than a cookie — the callback is a cross-site navigation and a SameSite=Strict cookie is not sent with it, so the alternative was weakening the cookie policy for every session to serve one flow.
  • RBAC — the five roles from SRS §2.3 over a single permission vocabulary, with object-level Device Group scoping for the group-restricted roles. Endpoints declare a permission, never a role list.
  • Credential vault — AES-256-GCM envelope encryption with per-record data keys wrapped by a pluggable master key, bound to the owning row so a ciphertext cannot be replayed into another record. Master-key rotation re-wraps without re-encrypting.
  • Audit log — append-only and hash-chained, enforced both by chain verification and by database triggers that reject UPDATE, DELETE and TRUNCATE outright.
  • Secret scrubbing — one central processor redacts secrets from every log line and audit record, including device-config idioms like snmp-server community X.
  • Quality gates — ruff, mypy --strict, pytest, bandit, pip-audit, eslint, tsc, vitest, Trivy and Gitleaks, all wired into CI.

Phase 1 — inventory, credentials and jobs

  • Read-only enforcement — the four-layer guard described above, 19 platform policies, and netsecops-cli audit-commands to print them for review.
  • Device sessions — adapters hold a guarded session, never a transport, so there is no unchecked path to a device. SSH with host-key pin-on-first-use, jump hosts and per-device timeouts.
  • Inventory — devices, hierarchical Device Groups (ltree), sites, tags, and CSV import with a dry-run preview that reports the offending line before anything is written.
  • Credential vault — typed credentials whose secret fields are sealed and whose unknown fields are rejected, so a password cannot land in searchable metadata. Device assignments override inherited group ones, and group credentials are inherited down the hierarchy.
  • Job engine — scope resolution, per-device outcomes with FR-COL-07 error classes, credential fallback, graceful cancel, re-run-failed, idempotency keys, and a WebSocket progress stream.
  • Scope enforcement — Device Group visibility applied in the query, not by the caller, so a group-scoped user cannot widen their reach.

Phase 2 — Cisco collection and drift

  • Cisco parsers — IOS/IOS-XE, NX-OS and ASA configurations become a vendor-neutral Normalised Config Model. Parsing is tolerant: an unrecognised stanza is kept in raw_unparsed and never fails a collection, and the percentage understood is stored on the snapshot so a degraded parse is visible rather than silently weakening checks.
  • Provenance on every value — each NCM field records the artefact and line range it came from, so a finding can show the operator their own configuration line instead of asserting a conclusion.
  • Absent is not false — a service the configuration never mentions stays null, which later reports as Not evaluated. Only an explicit no ip http server becomes false. Blurring the two produces confident, wrong findings.
  • Collection profiles — what each platform is asked for, as data. A test asserts every profile command already appears in the §8.2 allow-list, so a profile can never widen what NetSecOps may send to a device.
  • Redaction before storage — secrets are replaced with fingerprinted placeholders on every path that leaves the server. The unredacted original exists in one place, sealed, reachable by one endpoint that needs config:view_unredacted and writes an audit record before it answers.
  • The command log is readable, not merely kept — an Evidence panel on each device's configuration page lists every command a collection issued, in order, with the hash of each original response, whether it succeeded, and how long it took. The unredacted body is a separate per-artefact request rather than a toggle, because a toggle would mean the browser already held the secret. A partial collection says so at the top, since that is what makes the Not evaluated results below it explicable.
  • Snapshots and drift — identical configurations de-duplicate to one row, ignoring volatile lines like NVRAM timestamps and ntp clock-period. Pin a snapshot as the baseline and later collections that differ raise a drift finding with the diff attached, severity raised for security-relevant changes.
  • Diff, two ways — a unified and side-by-side text diff, plus a semantic diff over the NCM that says "management.services.telnet.enabled changed disabled → enabled" rather than leaving an operator to derive it from ±40 lines.
  • Offline configuration upload — assess an air-gapped or pre-onboarding device from an exported configuration file, through the same storage, parsing and drift path as a live collection.

Phase 3 — the check engine and findings

  • A check engine, and three rules it never breaks. Checks are YAML — id, severity, applicability, JMESPath logic over the NCM, remediation, framework mapping. A field the parser never found yields Not Evaluated and names the missing path, never a verdict derived from its absence. An empty list is a real answer. A broken check is an Error against that check alone, so one bad file cannot cost an assessment.
  • 66 checks, all applicable to Cisco IOS — 47 declarative, 14 Python for logic YAML cannot honestly express, 5 regex. Against the fixture corpus the hardened switch scores 56 pass / 1 high-severity fail and the weak one 41 fails; the ASA reports 39 Not Applicable rather than passing switch checks it was never subject to.
  • Remediation is text, and only text. There is no field in the schema that could be executed, and a test asserts none appears (SRS §8).
  • Policies grouping checks, assignable to device groups, with per-check severity overrides — the customisation that matters, because severity is contextual in a way a shipped library cannot know. The CIS Cisco IOS L1 pack ships with 44 checks and is installed idempotently, then never overwritten.
  • Custom checks written through the API against the same schema the loader uses — and refused if they declare Python logic, since accepting a function name from a web form would let a user invoke any registered callable.
  • Golden-config templates as a fourth logic type (FR-DRIFT-04): named blocks of lines that must be present, or must not be, matched literally or by pattern and optionally required to be adjacent. Every block is evaluated rather than stopping at the first miss, because a template exists to say how far a device is from the build — "does not match golden" sends somebody back for another pass after each fix. A template describes one estate's own standard, so none ships in the library; the console's draft editor seeds one, which is the only place anybody can write one. Three authoring-time refusals guard the failure mode that matters, a template that cannot fail: no blocks at all, a pattern that will not compile, and a line pasted out of a redacted configuration — the last matches nothing on any device for ever, and says so at no point.
  • Exceptions with a mandatory expiry. The check still runs and its result is still stored; only the finding is suppressed. Hiding the result would make the compliance figure a fiction, and an exception without an end date is an undocumented decision.
  • Findings with a lifecycle that reflects reality. Resolved is reachable only by the check passing on a later assessment — the API refuses to set it by hand, so the status stays a measurement rather than a claim. A problem that returns reopens the original finding instead of appearing as a first sighting.
  • Whether it is getting better (FR-FIND-05, FR-CHK-09). Every other figure the product shows is a level — how many are open now — and none of them can tell an estate that has sat at forty criticals for a year from one that was at four hundred in January. The dashboard draws first sightings against lasting resolutions over ninety days with the median time to resolve beside them, and a device's page draws its risk score as it has actually moved. That series was already there: risk_scores has kept a row per assessment since Phase 3, because — as the model says in as many words — a single current number cannot show a direction. Nothing had ever read it. What is deliberately absent is open findings by severity over time, which FR-FIND-05 also asks for and this schema cannot honestly supply: a finding is one row carrying its current status, and reopening clears resolved_at, so that curve would show every fixed-and-returned problem as open throughout — wrong in exactly the estates worth charting. The page says how many came back instead, which is the size of what the resolved series cannot see.
  • A risk score that is documented and explainable. Severity weights are widely spaced on purpose: under a linear scheme fourteen Low findings outrank one Critical. Device criticality multiplies rather than adds. Not Evaluated is reported as a separate coverage figure instead of being quietly counted as a pass.
  • A letter per appliance, and a closure priority — both derived, neither invented. A Risk Trends page grades every firewall, switch and router A to F, buckets the open findings P1 to P4, and draws the estate's score day by day. Nothing on it is a new measurement, which is the whole design: the letter is the stored risk score in a band, and the priority is the severity weight times the criticality multiplier that the score is already summed from — so the grid the page shows a reader is computed by the function that assigned their buckets rather than written out beside it. The estate series carries each device's last reading forward, because a risk score is a level that held until the next assessment replaced it, and it takes a baseline from before the window so an estate assessed quarterly does not chart as a flat nothing. A device nobody has assessed has no letter at all — not A, which would call it clean, and not F, which would call it broken.
  • Compliance pivoted by framework control, with the percentage computed over what was actually decided — Not Applicable and Not Evaluated are in neither half.

Phase 4 — firewalls and rulebase analysis

  • Three more vendors — PAN-OS, FortiOS and Check Point. Check Point splits in two: the policy lives on the management server and the gateway holds only Gaia, and neither can answer the other's questions, so they are separate platforms rather than one parser guessing which it was handed.

    The Check Point management path is not proven against a real server, and there is specific reason to think it did not work. Two defects were found in one day by reading rather than by running, and both are now fixed. The collector handed the parser a single response where it expected every response keyed by command, so a management server parsed to zero rules with no error raised. And the rulebase request sent neither the access layer (name) nor the package, which every one of Check Point's own published examples includes; both are now discovered from the server and the query is issued once per layer.

    It remains unproven against real equipment. The requests match Check Point's published examples, which is the most documentation can establish — everything here describing Check Point rulebase analysis should still be read as "against a rulebase we were given" until somebody runs it at a live management server.

  • Rulebase normalisation — PAN-OS security rules, FortiOS policies and Check Point access layers become one ordered rule model, with objects and groups resolved so that analysis compares addresses rather than names.

  • Relationship analysis (FR-FW-03) — shadowing, redundancy, correlation and generalisation between rules, plus the hygiene findings that matter in practice: any–any rules, rules that log nothing, rules with no security profile, and unused objects. Rule negation is handled rather than ignored, since a negated source inverts the meaning of every comparison downstream.

  • NAT analysis (FR-FW-04) and manager child enumeration — Panorama, FortiManager and Check Point SMS, behind an approval gate, because discovering devices through a manager adds targets that nobody explicitly onboarded.

  • A rulebase viewer (FR-FW-06, FR-FW-07) with rule query and CSV export, so a finding about rule 1,847 can be looked at rather than taken on trust.

  • 5,000 rules analysed in 2.7s against a two-minute budget (NFR-PERF-03) — 8.7M rule pairs considered, 9,453 fully compared. The rulebase is shaped like a real one, overlapping /24s drawn from a shared object pool rather than a corpus where nothing intersects and only the prefilter is exercised. A deliberately adversarial rulebase where almost every pair overlaps still completes in 38s.

Phase 5 — wireless and AAA

  • Wireless from three controllers into one vocabulary — Cisco WLC AireOS, Catalyst 9800 and FortiGate. AireOS is a command list rather than a configuration file and the other two are hierarchical, but an SSID accepting WPA2-PSK is the same finding on all three, so the security posture normalises even where the syntax cannot.
  • AAA servers as first-class targets — Cisco ISE and FortiAuthenticator over their REST APIs, FreeRADIUS and tac_plus by reading their configuration files over SSH. These fill aaa_server rather than aaa: they are the service the estate authenticates against, not a consumer of it.
  • Cross-estate correlation (FR-AAA-05) — every device's configured AAA servers against the servers in inventory, and every server's client list against the devices in inventory. The highest-value output is the second direction: a switch configured on ISE but absent from inventory is a device assessed by nothing, and a clean compliance percentage measured over an estate that does not contain it.
  • Three conclusions it refuses to draw. With no AAA server collected, every device trivially appears on no client list — reported as the absence of the question, never as "every device is unregistered". ISE and FortiAuthenticator mask shared secrets, so reuse is unknown for their clients rather than absent. Coverage over an estate nothing was collected from is null, never 0%.
  • An AAA posture dashboard (FR-AAA-06) — coverage, accepted protocols, orphaned clients and a certificate expiry timeline, each panel stating where it is blind. The protocols panel is titled "accepted", not "in use", because nothing here observes a live authentication. A certificate whose expiry could not be read is listed as undated rather than dropped: an unreadable date is not a distant one.
  • 13 wireless and AAA checks, and an audit that every check expression resolves against real parser output — a check naming an NCM path no parser populates is not a dead check but a false finding on every device, forever.
  • The Phase 5 acceptance criterion, as a test. Nine devices built from the shipped fixtures through the shipped parsers, with every expected number read off the fixtures by hand rather than off a run of the code (test_phase5_acceptance.py).

Phase 6 — vulnerability assessment

Vulnerability assessment produces findings now. A device is assessed against ingested advisories and end-of-life data, and the result is a finding on that device with its own lifecycle, visible in the console at /vulnerabilities.

It runs on two triggers, and both are needed because they answer different halves of one question. A collect-and-assess job weighs the catalogue against the configuration it just collected — this device changed, is it exposed? A vuln-rematch job, created by a feed import, re-weighs the estate against its stored snapshots without contacting anything — the catalogue changed, is anything newly exposed? An advisory published on Tuesday can make Monday's unmoved software exploitable, and finding that out should not require a collection window across five hundred devices.

The parts that shape the answer:

  • Four outcomes, not two. The matcher returns confirmed, likely, not affected or not evaluated, and the last is the default. A finding opens on confirmed or likely and is resolved only by a positive not-affected; not-evaluated leaves it open. Clearing a device requires evidence, never the absence of it.
  • Only a fully-understood advisory can clear a device. An advisory whose version ranges were partly unparseable can rule a device in and never out, which is why CSAF ingestion keeps what it cannot read instead of dropping or guessing it.
  • Offline bundle import, hash-verified. POST /vulnerabilities/feeds/import takes NVD 2.0 JSON, CSAF 2.0, endoflife.date, CISA KEV and FIRST EPSS bundles — the last as gzipped CSV, which is what FIRST actually publishes. SHA-256 is checked before anything is written, and every attempt is recorded including the failures.
  • KEV is prioritisation you can act on. Is this being exploited right now outranks every severity score: a CVSS 9.8 nobody has ever attacked and a 7.5 in active ransomware use are not the same work item. Importing the catalogue writes False onto every CVE it does not list, which is the whole difference between three states and two — set only the listed ones and everything else still reads "never checked", so the filter still matches nothing.
  • The catalogue is stored whole, not reduced to a flag. Otherwise the answer depends on import order: load the catalogue, then an advisory bundle introducing a new CVE, and that CVE reads "never checked" while an entry for it sits in the same database. With the catalogue present the flag is derivable whenever a CVE arrives, either way round.
  • EPSS scores what is known and leaves the rest null. A CVE the feed does not mention is unscored; writing zero would say "almost certainly not exploited", which for anything too new to have been modelled is exactly backwards.

The parts added last, each of which changes what an answer means:

  • Scheduled sync is built (FR-VUL-07). POST /vulnerabilities/feeds/sync queues a job, and the scheduler fires one nightly. Three design points carry it. The online route reuses the offline importer wholesale rather than growing its own ingest — otherwise the code air-gapped customers depend on is not the code anyone exercises daily, and only the run's recorded mode distinguishes them. NVD is fetched incrementally by last-modified date, because a CVE whose score or affected ranges were revised is exactly when a device's status changes without the device changing, and a "new CVEs only" design misses it. And offline mode refuses out loud: a scheduled sync that appears configured, never runs and reports no error is the stale-feed-that-looks-current failure this whole subsystem exists to prevent.

    A gap wider than NVD will answer — it caps a query at 120 days — comes back partial with the uncovered stretch named, not succeeded. A sync claiming currency over three months it never asked about is the same confident-wrong answer in a smaller box.

    Vendor PSIRT feeds are not fetched: Cisco's openVuln API needs an OAuth client credential and the others publish CSAF at per-advisory URLs that must be walked from an index. Both already have a working offline path.

  • CPE product names are verified against the NVD dictionary, and corroborated against evidence. The two answer different questions and both are needed. GET /vulnerabilities/cpe-coverage compares the platform-to-CPE table against the CPE strings imported advisories actually use, reporting each name as corroborated, contradicted (the same name under different punctuation appears instead) or no evidence. Deliberately not string similarity: ios_xe and ios_xr are 0.8 similar and are different operating systems.

    That method cannot find a name nothing has ever published, because absence is no evidence rather than a contradiction — so the table was also queried directly against the live NVD CPE dictionary on 2026-09-19. Eleven of thirteen matched, from 71 entries for fortinet:fortiauthenticator to 6,474 for cisco:ios. Two did not exist at all: checkpoint:security_management, for which NVD has no management-server product, and shrubbery:tac_plus, for which there is no such vendor. Both were matching nothing while looking like a clean result, and are now recorded in NO_DICTIONARY_ENTRY with the evidence rather than guessed at.

  • Chassis coverage is partial and bounded by what NVD names. Hardware advisories are matched on the model the device reports, which is an unbounded set rather than a table that can be verified in advance. Of eight real model strings checked, three resolved (C9300-48P, PA-3220, PA-850), two differ from NVD's spelling (FortiGate-100F against fortigate_100f, ASA5525 against asa_5525-x) and two have no NVD entry under any spelling. No single normalisation closes the gap — Palo Alto matches keeping its hyphen where Fortinet needs an underscore — so an unmatched chassis produces no finding rather than a wrong one, and the gap is recorded rather than papered over.

  • The upgrade-path view is built (FR-VUL-10). GET /vulnerabilities/devices/{id}/upgrade-path ranks every release the device's own advisories name as fixed by what each would close, KEV first — a release ending one vulnerability under active exploitation beats one ending nine nobody has attacked.

    Candidates are never synthesised. Suggesting "try 17.9.5" because 17.9.4 is fixed would recommend a release that may not exist, and an engineer who schedules an outage for it does not get a second one.

    Each CVE comes back eliminated, remaining or undetermined, and the third is never folded into the others: 15.2(7)E3 and 15.2(4)M5 are parallel trains with independent fix schedules, and neither is later than the other. Calling such a CVE fixed is dangerous; calling it still-open is safer and still wrong, because it makes a good upgrade look worse and steers the engineer toward a release that closes less.

    A CVE is only eliminated when every advisory naming it is closed. One flaw routinely appears in several — a vendor re-issues, or it affects two components with different fix trains — and an earlier version of this credited a release with a fix it only partly delivered.

The foundations underneath

  • Versions that are not a total order. packaging.Version and every semver library assume any two versions can be ranked. Cisco IOS breaks that: 15.2(7)E3 and 15.2(4)M5 are parallel trains with independent fix schedules and neither is later. Forcing an order there is not approximately right — it reports a patched device as exploitable, or an exploitable one as patched, depending which way the comparison falls. compare() returns "not ordered" as a third answer, and DeviceVersion has no __lt__, because < cannot express it.
  • Operational state reaches the parsers. Version, model and serial are not in a running configuration on most platforms; they come from show version and friends, which the collection runner used to store as an artefact and then discard. Supporting artefacts are now carried beside the configuration — never inside it, because show version reports an uptime that would make every device drift on every poll.
  • CPE 2.3 identifiers, with two refusals. No CPE without a version, because a wildcard matches every advisory ever written for the product. No CPE for an unmapped product, because a guessed name matches nothing while looking like a clean result. The product names are provisional until checked against a real NVD dictionary — unverified_products() exists for exactly that, and until it runs any wrong name is a device silently reporting zero vulnerabilities.
  • CSAF 2.0 ingestion that keeps what it cannot read. A prose version range or a product id the document never defines is recorded as unparsed with the vendor's own text, not dropped and not guessed. Dropping it hides a real vulnerability; guessing flags every device running the product. An advisory that is only partly understood can rule a device in, never out.

Phase 7 — discovery, reporting and integrations

Reporting is built, and reports are dated artefacts rather than saved queries. A report's content is assembled once, hashed, and never recomputed: re-reading March's report in September returns March's numbers, including findings that have since been fixed. That is what lets it answer "what did you know on 31 March", which no live view can. All nine catalogued templates assemble, in four formats — JSON, CSV, XLSX and PDF — and every format of one report carries the same content hash, because they render the same frozen content. A template with no single table is refused for CSV and XLSX rather than emitting a blank grid that reads as "no findings".

Discovery probes, and every run is paced. Scopes, the FR-DISC-02 probe allow-list, fingerprinting with confidence scoring and the pending-review queue were built first and had nothing driving them; FR-DISC-05 supplies the rest. A run is a job: it is queued, cancellable between batches, and recorded as a discovery_runs row that outlives the job history. All five permitted probes are sent — ICMP echo, TCP connect to the scope's ports, an SSH banner read, an HTTPS certificate-and-header fetch, and an SNMP read of sysObjectID and sysDescr.

The rate limit is the reason this could ship at all. An unpaced run across a scope is the port sweep SRS §1.2 forbids, whatever the allow-list says about the individual packets, so the prober holds the limiter and there is no code path from the endpoint to a socket that skips it. The default is FR-DISC-05's 50 hosts a second, configurable per scope up to a ceiling — "configurable" with no ceiling would make the requirement unenforceable.

SNMP is read, and its community is a credential. sysObjectID names the exact hardware model from a vendor-assigned tree where an SSH banner says "Cisco" at best, so a host that answers it usually needs no human at all. The community string is stored against the scope and sealed in the vault rather than kept in a column: public is a credential too, and trying it is a credential guess whatever its reputation. v2c only — SNMPv3's User Security Model needs a username and two keys per device, which nobody has for a host they have not yet identified.

Two things a run cannot do are recorded on the run itself rather than left to look like a quiet network. A scope that asks for SNMP without a usable credential still sends the other four probes and records a caveat; the cost is every host's fingerprint confidence, and a queue full of low-confidence entries otherwise looks like a hard-to-identify estate rather than a missing credential. ICMP needs CAP_NET_RAW, which containers withhold by default; without it liveness falls back to TCP and a device with no open port on the list is missed. Both appear beside the counters in the console, because "0 hosts found" and "0 hosts found, and nothing could be asked" are different answers.

Scheduling is built (FR-JOB-02, and FR-DISC-05's second half). netsecops-cli scheduler is a separate process that fires due schedules and enqueues them through the same path the API uses, so a scheduled collection and a manual one are the same job. Four behaviours are where the obvious implementation is the wrong one: a scheduler down for a day fires each schedule once, not once per missed occurrence; a blackout window skips rather than defers, because deferring stacks every skipped schedule onto one minute; two schedulers never fire the same schedule (FOR UPDATE SKIP LOCKED); and a schedule with no possible slot is disabled with a reason rather than silently never running. Cron is read in the schedule's own time zone — 0 2 * * * in Asia/Kolkata is not 02:00 UTC.

SIEM forwarding is built (FR-INT-02). Findings and audit records go to a collector as RFC 5424 syslog over TLS, in CEF or JSON. It is batch-and-watermark rather than send-on-write: emitting from every write path would put a network call inside the transaction that created the finding, so a dead collector would slow or fail the assessment that found it — inverting the priority, since the assessment is the product and the forwarding is a copy.

The two streams have different hazards. Audit records carry a monotonic id, so the watermark is exact. Findings are UUID-keyed, so theirs is a timestamp — and a row whose created_at is T can commit after a batch already advanced past T, and would then never be sent, with nothing anywhere looking wrong. The finding stream therefore stays thirty seconds behind the present, trading a little latency for not losing events silently. A failed send does not advance either watermark, so a collector outage delays delivery rather than dropping it.

Two details that are easy to get backwards and invisible when you do. Syslog severity runs 0 (emergency) to 7 (debug) — inverted relative to CEF's 0–10 — so a table written by analogy sends critical events as debug, where the first relay filtering on severity drops them while the integration looks healthy. And TLS framing is octet-counted, not newline-delimited, because a JSON body can legally contain a newline and framing on one splits records into fragments the collector cannot parse while the transport reports every byte delivered.

Nothing leaves unscrubbed: findings quote configuration, and configuration carries community strings and pre-shared keys, so the redaction that protects the database protects the wire too.

Notifications are built (FR-INT-01). E-mail over SMTP/TLS, HMAC-signed webhooks, and Slack and Teams incoming webhooks, with subscriptions filtering by event kind and severity — both, because "everything critical" and "every KEV match however it is scored" are different subscriptions a real operator wants.

Raising and sending are separate on purpose. raise_event writes a delivery row and returns; it runs inside whatever transaction produced the event, so a slow SMTP server cannot slow the assessment that found the problem and a dead one cannot roll it back. Sending happens in a worker job against a durable queue, because the notifications that matter most are raised when something is badly wrong — which is exactly when a process is most likely to be restarted.

Retry is bounded and ends somewhere visible. Retrying forever turns one receiver's outage into an unbounded queue; dropping after the last attempt loses the alert, which is worse. Attempts back off over roughly half an hour and then land in dead, kept and re-queueable.

Eight triggers, four call sites: four of them are already findings, so they are derived by scanning new findings and mapping the finding kind onto an event kind — which means a service that starts writing findings tomorrow gets notifications for free. Channel secrets follow the credential vault's contract exactly, and no read schema has a field through which one could come back out: a Slack incoming-webhook URL is a bearer credential.

Platform settings have a console (FR-ADM-01) at /settings, behind settings:read rather than any device permission — a channel's configuration decides where security alerts go, and a settings key decides how long evidence is kept. It lists deliveries that succeeded as well as those that failed, because "was anybody actually told?" is asked after an incident and a list of failures alone cannot answer it, and it surfaces any notification that gave up so an alert nobody received is visible rather than buried. The forwarding watermarks are shown but refused for write: editing one by hand silently skips or repeats part of the stream.

Scheduled report delivery is built (FR-RPT-04). A report is generated, frozen, and e-mailed as an attachment rather than a link — the people a compliance report is scheduled for are frequently the ones without a console login. The password-protected option is AES-256 and is refused rather than downgraded if the crypto backend is unavailable: the standard library can read an encrypted ZIP and cannot write one, so the obvious implementation ships an ordinary archive while reporting success, and the operator believes a document containing every finding in the estate is protected while it crosses two mail systems in the clear. Retention runs in the same job, and only deletes reports carrying an explicit expiry — the default for dated evidence is to keep it.

SNMP discovery is built (FR-DISC-02). sysObjectID is the heaviest fingerprint signal there is, so a host that answers it usually needs no human at all. The community lives in the vault as a credential attached to the scope, never in a column: public is a credential too, and trying it is a credential guess whatever its reputation — so the probe runs only where an operator supplied one, and a scope that asks for SNMP without a usable credential records a caveat rather than failing. SNMPv2c only; v3's User Security Model needs a username and two keys per device, which nobody has for a host they have not yet identified.

Still owed: ServiceNow/Jira (FR-INT-03, priority S) is not started. FR-INT-04, the RBAC'd REST API with OpenAPI, was already in place.

What is built

  • A discovery probe allow-list. SRS §1.2 rules out port sweeps, exploitation and brute-forcing; FR-DISC-02 names the five things discovery may do instead. That set is closed and asserted, because the way it erodes is not a bug but a drift — one more port for a customer running SSH on 2222, a slightly longer banner read, a second OID, each defensible alone and a port scanner in sum. A scope may name at most eight TCP ports, since "configurable list" otherwise permits a sweep assembled entirely from permitted probes. SNMP is refused outright without a configured credential: probing anyway means trying public, which is a credential guess. The SNMP GET is hand-written rather than pulled from a library — two OIDs from one request is about a hundred lines of BER, against a dependency carrying an async engine, a MIB compiler and a transport stack that none of this uses.
  • Scopes that refuse a mistyped prefix. 10.0.0.0/8 is one character from 10.0.0.0/18 and sixteen million probes from what the operator meant. The ceiling is counted from network sizes without expanding anything, exclusions are subtracted from the address space rather than filtered at probe time — so an excluded host is never enumerated at all — and the refusal names the likely cause, because an operator who reads only "over the limit" raises the limit.
  • Fingerprinting that scores what it could not tell apart. Signals are weighted by how much they actually prove — an SNMP sysObjectID far above an HTTP header — and confidence is capped below certainty, because no banner is proof. Conflicting signals subtract. The review queue then opens with the lowest confidence first, which looks backwards until you remember what the queue is for: a low score means the fingerprinter could not tell, and those are the entries that need a person.
  • A review queue that onboards nothing by itself. An approval carries the operator's corrections, a rejection carries a note, and neither deletes anything.
  • A paced run executor (FR-DISC-05). The limiter's unit is hosts, because the requirement's unit is hosts: one slot is taken when a host's probing begins, and that host's probes then run in sequence, so the packet rate stays proportional instead of multiplying by the port count. Concurrency is separate from the rate and answers a different question — how many hosts may be in flight while the slow ones time out — without which a scope of mostly-dead addresses runs at one host per timeout and the rate limit never binds at all. A cancel is honoured between batches, and the run keeps what it had already found rather than discarding it.
  • Managers as an inventory source (FR-DISC-06). Panorama, FortiManager and Check Point management enumerate their children. Preview and import are separate calls, children land in pending review rather than the inventory, and nothing is ever auto-deleted — a child that disappears from a manager may be a decommission or may be an API error, and the two must not be treated alike.
  • Reports as dated artefacts. Described above; the model and the freezing property live in db/models/reporting.py and services/reporting.py.

Phase 8 — topology and path analysis

Forwarding tables are data now, and they were not before. The NCM kept static_routes as an integer — a count — so the product could describe every rule on a firewall and had no idea which firewall a packet reached first. Each parser now reads its own platform's route grammar into a real list of destination, next hop, egress interface, protocol, distance, metric and VRF, and derives connected routes from interface addressing.

No new device access was needed for any of it. Static routes are in the running configuration, which was already collected and already parsed — the lines were being counted and thrown away, and on the ASA they sat on an explicit ignore list as noise. That is also how the commercial tools build their maps: no probing, no traceroute, no CDP/LLDP walk, no agents. SRS §8's read-only guarantee is untouched.

What a protocol learned is now collected too, which it was not at first. OSPF, BGP and EIGRP routes exist in no configuration file on any platform — they live only in the forwarding table — so the first version of this was static-and-connected only: complete for an edge or DMZ estate, partial in a routed core. Closing that needed show ip route (IOS), show ip route vrf all (NX-OS) and show route (ASA) added to the read-only allow-list, which is a change to what the product sends to a device and so was made deliberately and recorded in SRS §8.2 rather than slipped in. FortiOS already collected its table and nothing read it.

Every route still records the protocol that installed it, because the graph has to be able to say how complete it is rather than implying completeness. Three formats are parsed: IOS/ASA/FortiOS print a leading protocol code and a prefix, NX-OS prints a prefix line with indented *via lines and names protocols in words. Each has a way of failing silently — IOS prints subnetted children without a prefix length, so reading one literally gives a /32 host route to a network address that matches nothing; equal-cost paths arrive as continuation lines carrying no destination of their own; NX-OS ends its lines with the protocol and route type, so a reader working backwards adopts direct or a BGP tag as an interface name. All three are fixture-tested, and the operational table supersedes the configuration's statics rather than adding to them, since the device's own table already contains them.

Check Point now reads its forwarding table the way every other platform does, and the SNMP route walk that used to stand in for it has been removed. The removal is recorded rather than quietly done, because the way the walk went wrong is more instructive than the feature ever was.

The reasoning behind it was: five platforms reach their routing tables over the CLI, Check Point and PAN-OS have no route command in their collection profile, so read the table over SNMP instead. netsecops/snmp/ walked ipCidrRouteTable (RFC 2096) and merged what it found into the same routing.routes the parsers fill. It was carefully built — no encoder for a SET PDU existed at all, reject routes were excluded and counted, a truncated walk was recorded as truncated, and a failed walk cost a note rather than the snapshot. Every one of those is a good property of a thing that should not have existed.

Both halves of the premise were wrong, and neither was checked first.

The platforms do have route commands. Gaia's show route was already on the checkpoint_gaia allow-list in policies.py and simply never added to the collection profile — approved, permitted, never issued. PAN-OS answers <show><routing><route/></routing></show>, which the read-only guard already permits because it begins with <show>. The gap the feature existed to fill was a line of configuration in a profile.

And the MIB is not there anyway. Palo Alto's documentation states that PAN-OS "currently support[s] only the ipAddressTable and ipAddrTable in IP-MIB" — neither route table, in any version. Check Point does not document standard-MIB routing at all; sk90860 puts the routing table under the enterprise tree at .1.3.6.1.4.1.2620.1.6.6, so the standard OID the walk used was never the right one for Gaia either. Of the platforms checked, only IOS, IOS-XE and IOS-XR populate ipCidrRouteTable — and all three already collect routes over the CLI. FortiOS carries only the legacy ipRouteTable, ASA exposes a route count and nothing more, and NX-OS has neither. The comment calling it "the table every mainstream platform still populates" was an assertion, not a finding, and it was false.

The walk was also IPv4-only by construction (RFC 2096 types the index as IpAddress) and read only the default SNMP context, so per-VRF routes would have been silently absent — two more complete-looking partial answers in a feature written to prevent exactly that.

show route is now in the Gaia profile and parsed by parse_gaia_route_table. Gaia's legend reuses Cisco's letters for different things — its D is a BGP default where Cisco's is EIGRP, U is Unreachable where Cisco's is a per-user static, i is Inactive where Cisco's is IS-IS — so it has its own code map rather than sharing one, since the failure mode of sharing is a confident wrong protocol on a graph edge rather than a missing one. Routes the gateway is not forwarding on are dropped, and the interface is read by position: Gaia ends its lines cost 0, age 16426, and the shared helper that works backwards past anything resembling an uptime would return 16426 as an interface name.

PAN-OS is still outstanding. The op command is permitted and the response is XML, but its element structure is not documented — Palo Alto's own guidance is to read it off a live device or the API browser. Writing a parser against a guessed schema is how the Gaia password-policy parser came to match syntax Gaia never emits, with a fixture encoding the same fiction so the test agreed with the mistake. It waits for a real capture.

What was worth keeping is the BER codec unification: discovery and collection briefly shared one SNMP implementation instead of two. Collection no longer speaks SNMP at all, so discovery is again the only caller, but the codec stays in netsecops/snmp/ because it is a protocol rather than a discovery detail.

A path can now be traced across devices, and the answer has two axes. Give it a source, a destination, a protocol and a port, and it finds the devices in between and asks each of their rulebases — the FR-FW-06 rule query, run once per hop. Devices are joined by interface address: a route's next hop either is an address configured on another inventoried device or it is not, and matching on subnet instead would invent adjacencies on any shared transit link.

The two axes are the point, and they never collapse into one verdict:

Every firewall I found permits this, but I lost the path at 0.0.0.0/0 because 203.0.113.1 belongs to no device in the inventory.

That is routing: partially-routed, policy: partially-allowed — and allowed is only ever produced alongside routed. A permit speaks for the devices actually consulted, and an untraced remainder may hold another firewall; somebody opens a firewall on the strength of these answers. A block is the asymmetric case and stands on its own: the packet dies at the first denial, so what lies beyond it cannot change the result.

Three other distinctions the model keeps that a simpler one would lose. A router with no rulebase reports no decision rather than "allow", because a device that inspected nothing is not a control that was checked. A route in a VRF is never followed as if it were global — which VRF a packet is in depends on the ingress interface, and no parser records that binding, so the answer is unknown naming the VRF rather than unreachable. And a device whose snapshot predates route parsing, or whose table was truncated, cannot produce a negative answer at all.

The ranked missing-device report names the unmanaged next hops that terminate path analysis, ordered by how much reachability each conceals — a default route counts for far more than one specific prefix, and a next hop nine devices share counts for more than one. It is computed from the routes themselves rather than by running every path, which is quadratic in the estate and measures the queries somebody happened to ask instead of the gap itself. The addresses are evidence, not a work queue: an unmanaged next hop may be an ISP router, a customer handoff, or a virtual address no single box owns.

The graph is held between requests, and this reverses a decision the code used to state. services/topology.py refused to cache it — "a topology answer that silently reflects yesterday's estate is the kind of wrong that looks right" — and that objection is answered rather than overruled: every entry is keyed on a fingerprint of the devices and snapshots behind it, so a hit is byte-for-byte what a rebuild would produce and any change that could alter the graph forces one. It was worth doing because the graph costs ~580ms at 650 devices and every screen that touches topology built its own; a single dashboard load built two, in parallel, for one figure each. Measured on that estate, the dashboard pair went 1094ms → 313ms and the map 672ms → 63ms. Set NETSECOPS_TOPOLOGY_GRAPH_CACHE=false to go back to a rebuild per request.

The network map is the same graph before you know which question to ask: every device, drawn in columns by how many hops it sits from the estate's edge, with the strands between them. A strand exists only where one device's route names an address that is genuinely configured on another's interface — the join the path walk makes. Drawing a link because two devices share a subnet would be the tempting shortcut and would put wires on the picture that no packet follows; both ends of a transit /30 "contain" the next hop, and picking either invents an adjacency.

Three things it refuses to do. It does not omit its own boundary — an address the estate routes to and no inventoried device answers for is a box of its own, because that box names the device somebody would onboard to learn more. It does not fold a device that carries a rulebase into a bundle: fifty access switches collapse into one box with a count, which is the only reason a 650-device site is readable, and a firewall never does, because a control quietly taken off the picture is worse than a crowded one. And nothing about the layout is simulated — tier is breadth-first depth from the boundary, groups are the connected components of the graph, and every ordering falls back to the hostname, so the same estate draws the same way twice and this week's map can be compared with last week's.

Groups are structural and their names usually are not. Two devices are in one group when a packet can actually travel between them; the label is a site where one is configured and a shared hostname prefix otherwise, and the page says which — so nobody reads "london" as a site that somebody defined. Two sites defaulting to the same ISP address are deliberately not one group: no packet crosses between them.

Declared segmentation (FR-TOPO-07) is the same engine asked about every declared zone pair at once — the question people are actually audited on, which is not "what does this configuration do" but "is what it does what we said it would do". A zone is address space rather than a firewall's zone name, because dmz on one device and DMZ on another may be different things and a router has none at all. Intent is per ordered pair: most real segmentation is asymmetric, and a symmetric matrix would quietly assert the reverse of everything declared.

Each cell is evaluated by walking a packet, not by searching rulebases — a rule-centric check reports violations that do not exist (a permit on one firewall means nothing if a second denies it downstream) and misses violations that do (a permit assembled from one rule on the edge and another on the core belongs to no single rule to find). Not verified is never a pass: it has its own status, its own count beside the violations, and its own muted treatment, because a matrix showing green for pairs nobody could test is a compliance artefact asserting isolation that was never checked.

The page is a list rather than a grid, because a grid has a cell for every pair and must put something in the ones nobody declared — and whatever goes there reads as "fine". The policy is written from the same page. Withdrawing a statement sits inside the verdict it belongs to, where somebody is most tempted to make a red row go away, and says what it does: the requirement stops being made and the estate does not change. Removing a zone is refused while any intent names it — both foreign keys cascade, so the database would otherwise take a dozen requirements along with the tidy-up.

The matrix is also a report template, which closes a promise the code had been making and not keeping: services/segmentation.py refuses to store a verdict, on the grounds that an estate changes and a saved "compliant" is a claim about one that no longer exists — and points at the report archive, which had nowhere to put one. The frozen version cites the snapshots it was computed from, counts unverified pairs apart from both verdicts, and deliberately publishes no headline percentage: nine upheld out of ten is not ninety per cent segmented, because the tenth may be the pair that matters and the reader is meant to look at it.

The acceptance criterion is met (test_phase8_acceptance.py): a five-device fixture estate, a path query crossing three of them with the right traversed-device list and rule verdicts, and a query whose next hop belongs to no inventoried device naming that next hop instead of reporting it unreachable.

Writing it surfaced a contradiction in the specification. FR-TOPO-04 defines partially-routed as its own routing value; FR-TOPO-05 says the leaves-the-estate case is unknown. Both cannot hold — if that case were unknown, partially-routed would have nothing to describe. It is resolved in favour of the more precise value and recorded in SRS §3.8a: a path that leaves the estate at a named, real next hop is partially-routed, and unknown is kept for what genuinely could not be determined — a table never collected, a truncated one, a routing loop, a VRF binding nothing records. The difference is operational: the first is fixed by onboarding a device the report already ranks, the second by re-collecting one.

Still owed, and deliberately: IOS per-VRF tables. show ip route vrf <name> needs the VRF list first and a round trip per VRF, where NX-OS returns every table in one response — so IOS collects the global table only, and a path that depends on an IOS VRF resolves to unknown naming the VRF rather than guessing.

Two defects found by building the demonstration estate, both producing a confident blocked where the device forwards the traffic. That is the dangerous direction: a blocked verdict says a control is already in place, so somebody stops looking.

The first was in the reader. ip route 10.20.0.0/24 10.0.1.2 — NX-OS's prefix form — parsed its destination correctly and silently dropped the next hop, because the optional mask group matched 10.0.1.2 instead. Every NX-OS static route therefore pointed nowhere, the path walk could not follow it, and it reported the destination unreachable with nothing anywhere recording that a perfectly good next hop had been read and discarded. The existing test for that grammar asserted destinations only, which is how it survived: the assertion covered the half that worked.

The second was in the walk. A Cisco device's policy is a set of ACLs bound to different interfaces, not one ordered list. SecurityRule.rulebase records that and the hygiene analysis honours it — two entries in different ACLs are never compared, because they never see the same packet. The path walk was evaluating all of them as a single ordered list, so on a three-interface ASA the first ACL in the file decided every path; and since each ends in deny ip any any, every path came back blocked by a rule governing traffic in a different direction.

Both halves are now closed. An ACL bound to no interface — a vty filter, an SNMP filter, a leftover — filters nothing, and its rules are skipped. And the bindings are normalised, so the walk selects the list bound inbound on the interface the packet arrived on, plus any bound outbound on the interface it leaves by. Those are two separate first-match evaluations combined with deny-wins, never concatenated: merging them would let the first list's trailing deny shadow the second list's rules, which is the defect again in a smaller form. The binding shapes genuinely differ — IOS and NX-OS name a physical interface, an ASA names the nameif — so both names are kept and the match tries either, rather than resolving to one and losing the key that works elsewhere. A snapshot too old to carry bindings, or a device where no list names this hop's interfaces, still reports the decision as unknown rather than guessing.


What building it taught us

Three things went wrong in ways worth recording, because each one produced a confident answer rather than an error. They are the reason for several of the guardrails described above.

An entire rulebase matched nothing, and looked healthy

Fixing the selection meant the right access list was finally consulted — and it turned out nothing in it could match anything. Four spellings the ASA parser emitted, none of which the resolver reads:

Emitted Meaning Resolved as
ip every protocol nothing — unresolved
tcp/eq 443 log tcp port 443 nothing — unresolved
host 10.20.0.10 one address nothing — unresolved
subnet 10.20.0.0 255.255.255.0 (in an object) a /24 the empty set, silently

The first three at least landed in the rule's unresolved list. The fourth did not: an address object whose value could not be read returned an empty set rather than an error, so every rule referencing it was parsed, ordered, analysed — and able to match nothing at all, with no error anywhere. That is how an entire platform's rulebase can be inert while looking healthy.

The parser now emits the same vocabulary as the IOS reader, reusing that reader's port handling rather than keeping a second dialect — including its refusal to express neq, which marks a rule partial instead of quietly widening it. And the resolver now refuses an object it cannot read instead of returning nothing, so the next instance of this surfaces as an unresolved reference rather than as a rulebase that silently does nothing.

So every platform was swept

The mechanism was in the shared resolver, not in the ASA, so PAN-OS, FortiOS and Check Point were equally exposed and had never been checked. test_silent_emptiness.py now parses every device fixture in the repository, resolves its rulebase and fails on anything that is silently nothing: a rule that can never match, an object that stands for nothing, a configuration that parses to an almost-empty NCM. It also asserts that every platform with a parser has a fixture, because a platform nothing exercises is one whose rulebase can go inert unnoticed.

The rulebases are clean. That was the question worth asking and the answer is good.

What the sweep did find was one level up: a platform means three different things and they were being conflated. Three registries key off a platform name — POLICIES (what may be sent), PROFILES (what is sent) and PARSERS (how the answer is read) — and Device.effective_platform, documented as "the policy key to enforce against", was used as all three plus the check-applicability platform.

Enabling allow_expert on a Check Point gateway is a documented, supported act (ADR-002). Doing it broke the device twice: no parser is registered for checkpoint_gaia_expert, so it could no longer store a snapshot; and because applicability is an exact match against platforms: [checkpoint_gaia], every Check Point check reported Not Applicable and the gateway was assessed to nothing. Existing findings stayed open — only a PASS resolves one — but nothing new was ever evaluated, and a gateway that is not being assessed looks exactly like a gateway with nothing wrong. It is now policy_platform, and a rename rather than a fix in place because the old name is what invited the misuse.

Separately, five platforms could be assigned to a device and never collected from. _validate_platform checked only for a read-only policy, the narrowest of the three registries, while its own docstring said catching this at creation "is far kinder than failing mid-collection". It now requires a collection profile or a manager enumerator.

Address translation is declared, not modelled, and for the same kind of reason. A trace that crossed a translating firewall and reached the far side used to report routed / allowed, when the devices after that firewall had been asked about the addresses in the query rather than the ones the packet was carrying. Deciding whether a given packet is translated means reading each NAT rule's original and translated addresses, and those fields do not mean the same thing across the four parsers: PAN-OS puts the rule's source members in original whatever it translates, FortiOS puts a VIP's external address there, Check Point joins several originals into one string, and Cisco ASA sets neither — only the raw line and an interface pair. A matcher built on that would be wrong differently on each platform, and a wrong path verdict is somebody opening a firewall on the strength of it. So a path that continues past a device carrying NAT rules reports partially-allowed with the device named, and says that translation was not evaluated. A translating firewall at the end of the path changes nothing: the trace is over and the addresses it reasoned about were the ones asked for. Normalising NAT across the parsers is the prerequisite and is its own piece of work.

Twenty-six screens that tell you which one you are on

The console is plain CSS with custom properties — no utility framework, no UI kit (SRS §2.2) — and one consequence of one layout, one type scale and one accent across twenty-six screens was that the fastest way to know where you were was to read the heading. Each navigation group now has a colour, carried on its sidebar heading and on a rule under the page title of every screen it leads to. That mapping is derived from the navigation rather than passed to each page, and sections.test.ts walks the sidebar and fails if the two disagree — which they did on the first attempt, because Firewall Analysis reads as a Risk page and is filed under Network.

The palette deliberately avoids red, amber and green. The severity ramp is the only colour in this product that means anything, and a page header that could be mistaken for a verdict about what is on the page would be worse than a grey one.

Three layout faults went with it, all invisible until content got long. Table cells were white-space: nowrap, which does not truncate a finding's title — it widens the column, so one sentence pushed every other column off the side. Grid tracks written 1fr are really minmax(auto, 1fr) and refuse to shrink below their content, so a 64-character digest overflowed its column and sat on top of the next; every one is now minmax(0, 1fr), including the main content column, which had been letting a wide table stretch the whole page. And the split-diff header had two tracks where its rows had four, so neither label sat over the pane it named.

The stylesheet also had four lines of English prose being parsed as CSS — a comment had lost its opening delimiter, esbuild recovered, and the only trace was a build warning nobody read. stylesheet.test.ts now fails on unbalanced comment delimiters, stray prose and bare 1fr tracks.

The console can reach the whole product

Unreachable capability is indistinguishable from absent capability, and this codebase had already paid for that once: the vulnerability engine was built, unit-tested, and wired only to a job type nothing created, so for three phases it never ran and nobody noticed — "assessed and found nothing" and "never assessed" look identical on a dashboard.

So the API is audited against the console. scripts/api_reachability.py reconciles every published operation against every path literal in frontend/src, and docs/api-reachability.md records the result and the triage. The first run found 61 of 139 operations the console never requested, 32 of them in areas with no page at all — no way to create a user, define a policy, browse the check library, schedule an assessment, issue an API token or file a risk-acceptance exception without a REST client.

Those 32 are now closed, along with the credential vault, job cancellation, feed import and the raw-evidence reads before them. What remains is 16: 14 operations on pages that exist and do not yet call them, plus GET /metrics and GET /readyz, which are correctly machine-only.

Building the surfaces found four defects that no test had caught, each invisible for the same reason — the wrong behaviour and the right one produced identical output while nothing exercised the path:

  • DELETE /credentials/assignments/{id} took an id nothing but the creating POST ever emitted, so a credential binding made last month could not be withdrawn from any client. Granting access is recoverable; being unable to withdraw it is not.
  • PUT /users/{id}/scope wrote a Device Group scope no response carried, so an administrator could set one and never read it back.
  • POST /sites accepted a location, returned one, had a column for it, and never stored it — every site read back null, which looks exactly like a field nobody filled in.
  • GET /checks/{id} returns a check's expression but not its applicability or assertion, so a shipped check cannot be copied as the starting point for a custom one. Left open and recorded: closing it is an API change, not a page.

Treat a rise in the unreferenced count on a pull request the way you would treat a drop in coverage.

Repository layout

netsecops/
├─ backend/
│  ├─ netsecops/
│  │  ├─ adapters/     # read-only guard, platform policies, sessions, transports
│  │  ├─ api/          # FastAPI routers, dependencies, middleware
│  │  ├─ core/         # config, logging, crypto, security, RBAC, errors
│  │  ├─ db/           # declarative base, session, models, Alembic migrations
│  │  ├─ checks/       # the check engine, its YAML library and policy packs
│  │  ├─ discovery/    # probe allow-list, scopes, fingerprinting, pacing, the
│  │  │                #   probe transport and the run executor (Phase 7)
│  │  ├─ topology/     # the layer-3 graph, the path walk and the ranked
│  │  │                #   missing-device report (Phase 8)
│  │  ├─ firewall/     # rulebase model, relationship analysis, NAT, hygiene
│  │  ├─ ncm/          # the Normalised Config Model (NCM v1)
│  │  ├─ parsers/      # vendor config parsers, one package per vendor
│  │  ├─ schemas/      # Pydantic request/response models
│  │  ├─ services/     # business logic, independent of HTTP
│  │  ├─ vuln/         # versions, CPEs, advisories, feed parsing (Phase 6)
│  │  ├─ workers/      # job runner, credential probe, queue abstraction
│  │  └─ cli.py        # netsecops-cli
│  └─ tests/
│     └─ fixtures/     # anonymised configs and operational output, by platform
├─ frontend/           # Vite + React 18 + TypeScript SPA
├─ scripts/            # smoke_test.py — post-deployment verification
│                      #   api_reachability.py — the console-vs-API audit
├─ tools/              # build_brand_assets.py — derives the served brand assets
├─ deploy/             # Dockerfiles, docker-compose, Caddy, Postgres init
└─ docs/
   ├─ brand/           # the master logo artwork, committed unmodified
   └─ ...              # SRS, ADRs, device-account guidance, deployment

Vendor-specific logic stays inside adapters/, parsers/ and the vendor packs under checks/library/; core services remain vendor-agnostic (C-6).

The files worth reading first

File Why
adapters/readonly.py The four-layer guard that enforces SRS §8
adapters/policies.py Exactly what NetSecOps may send to each platform
adapters/session.py Why there is no unchecked path to a device
adapters/profiles.py What each platform is actually asked for, and why
ncm/models.py The vendor-neutral model every check reads
services/snapshots.py How a change is told apart from noise
tests/test_readonly.py 283 assertions that the guard decides correctly
tests/test_device_session.py That nothing else reaches a real SSH server
tests/test_profiles.py That a profile cannot widen the device-facing surface
discovery/probes.py The five things discovery may send, and why the list is closed
vuln/versions.py Why two versions are sometimes not ordered at all
checks/engine.py Why a check declines to have an opinion
checks/library/ Every check, as data, with its reasoning
services/risk.py The risk formula, and why it is shaped that way

Development

make check       # every gate: lint, typecheck, test, security
make lint        # ruff + eslint
make typecheck   # mypy --strict + tsc
make test        # pytest + vitest
make test-cov    # backend coverage report
make security    # bandit, pip-audit, npm audit

Backend tests run against a real PostgreSQL instance (the schema uses JSONB, INET and advisory locks, so a substitute engine would not test what ships). make db starts one; override the target with TEST_DATABASE_URL.

Load tests are opt-in, because they commit real rows and take seconds rather than milliseconds:

cd backend && ../.venv/bin/python -m pytest -m performance -s

They measure job-queue throughput against NFR-PERF-01 — 20 workers sustain roughly 100× the required device-claim rate, which is the evidence behind ADR-001.

Conventions

  • Python 3.12+, type hints everywhere, mypy --strict clean.
  • Every schema change is an Alembic migration — no manual DDL. CI runs alembic check to catch a model that drifted from its migration.
  • Every endpoint appears in the authorization matrix (tests/test_authz_matrix.py). Adding a route without an entry fails the build by design.
  • All configuration via environment variables; no secrets in the repo or an image.

The command line

netsecops-cli create-admin           # bootstrap the first Super Admin
netsecops-cli generate-master-key    # generate a credential-vault master key
netsecops-cli generate-secret-key    # generate a JWT signing key
netsecops-cli verify-audit-chain     # replay the hash chain, detect tampering
netsecops-cli rotate-master-key      # re-wrap every stored secret
netsecops-cli reset-password <user>  # break-glass password reset
netsecops-cli reset-mfa <user>       # break-glass: clear a lost authenticator
netsecops-cli audit-commands         # print the read-only allow-list per platform
netsecops-cli permissions            # print the role × permission matrix
netsecops-cli health-check           # database reachability + schema revision
netsecops-cli show-config            # effective configuration, secrets masked
netsecops-cli demo-seed              # a demonstration estate, no device contacted
netsecops-cli demo-purge             # remove it, leaving any real device alone
netsecops-cli version                # build version

Two commands are not utilities but long-running processes:

netsecops-cli scheduler              # fire due schedules (FR-JOB-02)
netsecops-cli worker                 # execute queued jobs (FR-JOB-05)

Both run as their own containers under the workers compose profile, and running several of either is safe — schedules and jobs are both claimed with FOR UPDATE SKIP LOCKED, so a second process passes over what the first holds.

They are a pair, and the pairing is easy to miss. The scheduler creates job rows; it does not run them. Without a worker, a scheduled collection, feed sync, notification dispatch or report sits queued for ever — which the console shows as a job that never starts rather than as an error.

Locked out of MFA?

Disabling MFA through the API needs a session, and MFA is what blocks sign-in — so recovery runs on the server:

docker compose -f deploy/docker-compose.yml exec api netsecops-cli reset-mfa admin

The user then signs in with their password alone and re-enrols from Profile → Two-factor authentication. The reset is recorded in the audit log. Recovery codes issued at enrolment also work as a second factor — each one once.


Smoke-testing a deployment

python scripts/smoke_test.py \
    --api http://localhost:8000 --ui http://localhost:8080 \
    --admin-user admin --admin-password '...'

Checks the live stack across fourteen sections: unauthenticated rejection, security headers, cookie flags, RBAC, MFA enrolment and challenge, refresh rotation, audit-chain integrity, inventory and the SPA. It creates a throwaway user for the destructive parts and deletes it afterwards, so it never alters the account it signs in with.


Documentation

Document Contents
docs/SRS.md Full software requirements specification — the baseline
docs/device-accounts.md Recommended read-only accounts per platform
docs/evaluating.md Ten minutes with no device: the demonstration estate, what to look at, and the known limits
docs/commercial.md Sizing, what the incumbents charge, and what follows
docs/deployment.md Deployment, sizing, backup and key management
docs/api-reachability.md Which API operations the console can reach, and which need a surface
docs/parser-validation.md Measuring the parsers against configurations we did not write, and what that found
docs/vendor-research.md What Cisco, Palo Alto, Fortinet and Check Point's own documentation says we are missing
docs/adr/ Architecture decision records

The API documents itself: OpenAPI at /api/v1/openapi.json, interactive docs at /api/v1/docs outside production.

The console reaches every area of it. 151 operations are published and the console requests 140; the 11 it does not are individual operations on pages that already exist, plus the two probe endpoints, and each is named with what its absence costs. That is measured rather than estimated (scripts/api_reachability.py) and tracked in docs/api-reachability.md, because capability nobody can reach is indistinguishable from capability that does not exist. The measure earns its keep: it caught the segmentation page evaluating a policy that could only be written through the API, which made the feature read as empty on any console-only deployment.


Licence

MIT — see LICENSE.

About

A Network Security Posture Management (NSPM) tool for network and network security engineers to have a continuous view of the risks within their network.

Topics

Resources

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages