See the blast radius of a compromised dependency, powered by HydraDB.
actik is a graph-native software supply-chain intelligence platform. It models packages, versions, dependencies, vulnerabilities, maintainers and the internal applications that consume them as a single graph in HydraDB, then answers the questions that matter when a package is compromised:
- Which internal services are transitively exposed?
- Which version introduced the vulnerability? (and which version fixes it)
- Which applications resolved the bad version while it was live?
- Which other packages share a maintainer or infrastructure with it?
- Are there likely typosquat packages nearby?
- What is the complete blast radius?
Built for the Hack Hydra Track 02A brief (Repos, Dependencies + Code as Graphs) around the HydraDB open-source repo.
Supply-chain attacks via npm and PyPI are surging, and current developer tools fail to give real-time, deep context on malicious dependencies. This is fundamentally a graph traversal problem, not a semantic similarity problem. When a package is compromised, a defender needs the answer to a set of relationship questions in seconds:
- Which internal services are transitively exposed?
- Which version introduced the vulnerability, and which version fixes it?
- Which apps resolved the bad version while it was live?
- Which packages share maintainers or infrastructure?
- Are there likely typosquat packages nearby?
- What is the complete blast radius?
The TanStack compromise illustrates the stakes: 84 malicious artifacts across 42 packages in six minutes, self-propagating into .claude/ and .vscode/ in a way that survived npm uninstall. The defender's problem is speed — when a package is compromised at 09:00, which services are exposed by 09:06? That is a transitive reverse-dependency closure over a versioned ecosystem graph. A vector index cannot answer it at all, and a relational store answers it with awkward recursive CTEs. A graph database answers it natively.
actik models the whole supply chain as a graph in HydraDB and answers every one of those questions by traversal.
- Hono + TypeScript backend running on Bun
- HydraDB
graph-nodealongside it in Docker Compose (ghcr.io/hydra-db/hydradb) - Real graph traversals, blast radius is
algo.SSPathsoverDEPENDS_ONedges - Repo scanning with exact, lockfile-grounded version resolution
- Time-travel exposure windows (the track's 09:00 compromised, 09:06 exposed scenario)
- Worm and propagation simulation plus a live OSV watch loop
- 140+ automated tests
flowchart TB
subgraph internet["Internet"]
Web["actik Web (Next.js on Vercel)"]
end
subgraph vps["OVH VPS, Docker Compose"]
nginx["nginx (HTTPS, 80/443 only)"]
api["actik API (Hono + Bun, :8000)"]
hydradb["HydraDB graph-node (:8443 / bolt:7687 / :9090)"]
certbot["certbot (Let's Encrypt)"]
nginx --> api
api --> hydradb
certbot --> nginx
end
subgraph data["Data sources (polled during ingestion, not per request)"]
npm["npm Registry + bulk audit"]
pypi["PyPI JSON API"]
osv["Google OSV"]
end
subgraph scan_targets["Scan targets (on demand)"]
gh["GitHub / GitLab"]
end
Web -- HTTPS --> nginx
api -- ingestion --> npm
api -- ingestion --> pypi
api -- ingestion --> osv
api -- scan --> gh
Only nginx publishes ports (80/443). The API and HydraDB are reachable only on the internal compose network, so HydraDB stays private behind the API.
actik's whole pitch is graph-native supply-chain defense. The supply chain is relationship-heavy data, a package is meaningless in isolation. The important questions are all traversals:
A depends on B
B depends on C
C is vulnerable
D depends on A
E shares a maintainer with B
Every feature in this repo leans on a single graph stored in HydraDB:
- Blast radius is
algo.SSPathsoverDEPENDS_ONedges, a traversal a relational DB answers with recursive CTEs and a vector DB cannot answer at all. - The repo scanner writes
Repository - HAS_LOCKFILE - Lockfile - RESOLVES - PackageVersioninto the graph, then answers which resolved versions have advisories with a singleAFFECTED_BYjoin. - The exposure score re-traverses the same graph after each proposed fix to prove the fix clears the blast radius.
- Time travel answers which applications resolved the compromised version while it was live by reading
scanned_atonRESOLVESedges andpublished_at/modified_aton advisories, a temporal predicate over graph edges no other store in this problem has the shape for.
Everything the API returns (paths, chains, exposure windows) is produced by traversing the graph, not by querying a table.
erDiagram
Package ||--o{ PackageVersion : HAS_VERSION
Package ||--o{ Maintainer : MAINTAINED_BY
PackageVersion ||--o{ Advisory : AFFECTED_BY
PackageVersion ||--o{ PackageVersion : DEPENDS_ON
Repository ||--o{ Lockfile : HAS_LOCKFILE
Lockfile ||--o{ PackageVersion : RESOLVES
Alert ||--o{ PackageVersion : ALERTS_ON
Alert ||--o{ Lockfile : EXPOSES
Organization ||--o{ Repository : CONTAINS
The critical relationship is A DEPENDS_ON B, version-to-version exact. Each dependency declaration resolves to a single target version, preferring the version a real lockfile resolved for the same source (lockfile-grounded) and falling back to the best range match. This keeps blast radius trustworthy, so express@5.2.1 points at qs@6.5.2, not at every qs release.
Advisories carry both ends of the vulnerability story:
fixed_versions, the first version that fixes the advisory.introduced_versions, the first version that was vulnerable.
Together they answer which version introduced it and which version fixes it for every affected package.
- Package overview, versions, dependencies, dependents and maintainers
- Advisory details with affected versions, introduced and fixed versions
- Shared-maintainer analysis, packages connected through the same maintainer
- Typosquat candidates, similar package names with a risk score and the reasons (Levenshtein distance, character-substitution, scoped vs unscoped, popularity)
The primary feature. Given lodash@4.17.20, actik runs a reverse DEPENDS_ON traversal inside HydraDB and returns:
- direct and transitive dependents
- maximum dependency depth
- every affected repository, with the exact resolved version, the requested range, the internal
node_modulespath, and the full dependency chain from the app down to the compromised package - affected applications (repos whose lockfile kind is an application)
- traversal latency
flowchart LR
lodash["lodash@4.17.20"] --> a["package-a"] --> s1["service-a"]
lodash --> b["package-b"] --> s2["service-b"]
lodash --> c["package-c"]
s1 --> prod["production"]
s2 --> prod
style lodash fill:#ef4444,color:#fff
style prod fill:#6366f1,color:#fff
POST /api/scan accepts a repository URL or owner/name and pulls package-lock.json, yarn.lock, pnpm-lock.yaml, bun.lock, uv.lock or requirements*.txt straight from the repo with no clone and no token on GitHub or GitLab. It returns:
- an exposure score (0 to 100) weighted by severity times count
- every vulnerable package with its advisory and the exact fix command
- the exposure path from the app through the dependency chain
- a minimal-fix set, the fewest upgrades that clear every finding, each one verified by re-traversing HydraDB (the target version must exist in the graph and resolve to zero advisories)
Findings merge graph-backed AFFECTED_BY edges with a live Google OSV check, so a scan is useful on any lockfile, even packages never ingested. Unlinked packages are reported transparently.
GET /api/advisories/:id/exposure-window partitions the apps resolving an affected version into:
EXPOSED, the app scan happened during the advisory live window.AT_RISK, the app still resolves an affected version but was scanned outside the window.NOT_AFFECTED, the app resolved the package but a version outside the affected range.
Pass ?asOf=YYYY-MM-DD to replay the graph as it was at that date.
GET /api/simulate/propagation/lodash/4.17.20?compromisedAt=2026-05-14T09:00:00Z&perHopMs=360000 compromises a package at time zero and walks the reverse DEPENDS_ON closure, computing each app time-to-exposure as exposedAt = compromisedAt + depth x perHopMs, the track's 09:00 compromised, 09:06 exposed cadence by default.
POST /api/watch/run polls Google OSV for every version a scanned app resolves and records newly-flagged advisories as Alert nodes with a first_seen_at, linked to the affected version and every app that resolves it. GET /api/watch/incidents lists them newest-first with their exposure path.
Requires Bun.
bun install
bun run devOpen http://localhost:8000/health
Pulls HydraDB's published image from ghcr.io/hydra-db/hydradb.
cp .env.example .env
docker compose -f compose.dev.yml up --build| Service | URL |
|---|---|
| API | http://localhost:8000/health |
| HydraDB HTTP | http://localhost:8443 |
| HydraDB Bolt | bolt://localhost:7687 |
| HydraDB readiness | http://localhost:9090/readyz |
HydraDB data (store, cache, auth token) persists in ./hydradb-data; the compose entrypoint creates it on first start, so there is no manual setup.
cp .env.example .env
# set HYDRADB_AUTH_TOKEN to a real value (>= 32 chars)
# set FRONTEND_ORIGIN to the frontend's public origin (e.g. https://actik.xyz)
docker compose -f compose.prod.yml up -d --buildThat single command is the whole deploy:
- nginx starts with a throwaway self-signed cert (baked into its entrypoint) so it can serve the ACME challenge before Let's Encrypt has issued anything.
- certbot runs right after and issues a real Let's Encrypt certificate for your API hostname (default
api.actik.xyz, override withCERTBOT_EMAIL). - nginx detects the cert change and reloads itself, no manual step.
Only nginx publishes ports (80/443). The API and HydraDB are reachable only on the internal compose network, and nginx proxies https://api.actik.xyz to http://api:8000. HydraDB (8443/7687/9090) stays internal.
Nginx config lives in proxy/nginx.conf and the entrypoint in compose.prod.yml; change the API hostname in both if yours differs. Certs are valid 90 days and auto-renewed each up because certbot uses --keep-until-expiring.
Stop with docker compose -f compose.prod.yml down. Data lives in ./hydradb-data; back it up or mount it from a persistent volume.
Seeds a supply-chain graph into HydraDB from three sources:
- npm registry plus npm bulk audit advisories
- PyPI JSON API vulnerabilities
- Google OSV (advisories, affected ranges, introduced and fixed versions)
After the package graph is written, the runner ingests a synthetic organization from demo-org/ (see DEMO_ORG_PATH): each repository's lockfiles are parsed into Repository - HAS_LOCKFILE - Lockfile - RESOLVES - PackageVersion edges, keeping the exact resolved version, the requested range, and the internal node_modules path as evidence.
flowchart LR
subgraph sources["Data sources"]
npm["npm Registry + bulk audit"]
pypi["PyPI JSON API"]
osv["Google OSV"]
end
subgraph pipeline["Ingestion (separate from request path)"]
fetch["Fetch metadata"]
norm["Normalize (stable ids, range parsing)"]
write["GraphWriter (MERGE, idempotent)"]
end
demo["demo-org/ lockfiles"]
npm --> fetch
pypi --> fetch
osv --> fetch
fetch --> norm --> write
demo --> norm
write --> hydradb["HydraDB"]
The API auto-seeds on startup: if HydraDB has an empty package graph it runs the full ingestion pipeline before serving (fast restarts skip it when already seeded). No manual step is needed, but you can refresh manually:
docker compose exec api bun run ingestAll limits come from env vars in .env.example (INGESTION_MAX_PACKAGES, INGESTION_MAX_DEPTH, INGESTION_MAX_ADVISORIES, INGESTION_CONCURRENCY). Re-running is idempotent (stable vertex ids + MERGE).
Mounts at /api. Error responses use {"error":{"code","message"}}; 404 for unknown packages or advisories, 400 for invalid names, 429 when rate-limited.
All :name / :name/:version routes accept an optional ?ecosystem=npm|PyPI query parameter to disambiguate packages that share a name across ecosystems.
| Route | Purpose |
|---|---|
GET /api/packages/:name |
Package overview + version list |
GET /api/packages/:name/maintainers |
Maintainers of a package |
GET /api/packages/:name/shared-maintainers |
Packages sharing a maintainer |
GET /api/packages/:name/typosquats |
Similar-name candidates (npm search + Levenshtein) |
GET /api/packages/:name/:version |
Version details + its advisories |
GET /api/packages/:name/:version/dependencies |
Forward dependencies |
GET /api/packages/:name/:version/dependents |
Direct dependents |
GET /api/packages/:name/:version/blast-radius |
Direct/transitive dependents, max depth, paths, per-repository exposure paths, latency, affected repositories + resolution evidence |
GET /api/packages/:name/:version/graph |
Dependency neighborhood for the frontend graph view, incl. resolving repositories |
GET /api/advisories/:id |
Advisory details + affected versions + known introduced and fixed versions |
GET /api/advisories/:id/exposure-window |
Apps that resolved an affected version while the advisory was live with a per-app EXPOSED / AT_RISK conclusion; optional ?asOf=YYYY-MM-DD snapshot |
GET /api/graph/:name/:version |
Alias for the package graph endpoint |
POST /api/scan |
Scan a repository ({"repo":"owner/name"} or a full GitHub/GitLab URL): fetch manifests, resolve exact versions, write into the graph, return exposure score + findings + fixes + minimal-fix set |
GET /api/scan/:owner/:name |
Re-run analysis for a previously scanned repo from the graph (no re-fetch) |
GET /api/simulate/propagation/:name/:version |
Worm simulation: compromise a package at ?compromisedAt=, compute each app time-to-exposure from DEPENDS_ON depth (?perHopMs=, default 6 min) |
GET /api/investigate/:ecosystem/:name/:version |
One-call investigation: version details + advisories + blast radius + maintainer risk + typosquats + recommendations |
POST /api/watch/run |
Live-watch pass: poll OSV for every scanned/resolved version, record newly-flagged advisories as Alert nodes |
GET /api/watch/status |
Last live-watch run summary |
GET /api/watch/incidents |
Recent incidents, each with its exposure path (repo -> lockfile -> pkg@version -> advisory) |
Example, blast radius:
GET /api/packages/lodash/4.17.20/blast-radius{
"package": { "name": "lodash", "version": "4.17.20" },
"directDependents": 0,
"transitiveDependents": 0,
"maxDepth": 0,
"affectedRepositories": ["payments-api", "storefront"],
"affectedApplications": 2,
"repositoryPaths": [
{
"repository": "payments-api",
"lockfile": "payments-api/package-lock.json",
"internalPath": "node_modules/lodash",
"path": ["4.17.20"],
"depth": 0
}
],
"latencyMs": 3
}Example, advisory with introduced and fixed versions:
GET /api/advisories/GHSA-g4mx-q9vg-27p4{
"id": "GHSA-g4mx-q9vg-27p4",
"severity": "MODERATE",
"summary": "urllib3's request body not stripped after redirect from 303 status changes request method to GET",
"publishedAt": "2023-10-17T20:15:25Z",
"modifiedAt": "2026-02-04T03:30:16.767903Z",
"fixedVersions": { "urllib3": "1.26.18" },
"introducedVersions": { "urllib3": "1.20" },
"affectedVersions": [{ "name": "urllib3", "version": "1.26.4", "ecosystem": "PyPI" }]
}actik-backend/
├── src/
│ ├── index.ts # app wiring, CORS, rate limiting, error handling
│ ├── routes/ # HTTP routes (packages, advisories, scan, watch)
│ ├── services/ # business logic (blast radius, advisory, investigate)
│ ├── hydra/ # HydraDB client, Cypher queries, schema
│ ├── ingestion/ # npm/PyPI/OSV ingestion, normalization, graph writer
│ ├── analysis/ # blast radius, paths, maintainers, typosquats, propagation
│ └── lib/ # config, errors, logging, rate limiting, VCS clients
├── demo-org/ # synthetic org (repos + lockfiles) for the demo dataset
├── proxy/ # nginx config for production HTTPS
├── tests/ # 140+ unit tests
├── compose.dev.yml # dev docker compose (API + HydraDB)
├── compose.prod.yml # prod docker compose (nginx + certbot + API + HydraDB)
├── .env.example # all configuration knobs
└── README.md
bun test
bunx tsc --noEmitReal numbers, straight from HydraDB — nothing is invented. Requires a running, seeded HydraDB (the dev compose brings it up and seeds it automatically).
bun run bench # dataset report + query latency (P50/P95/P99)
bun run bench:all # same, plus precision/recall vs OSV ground truth
bun run bench --iter 100 # raise the iteration count for tighter percentilesWhat it measures:
- Dataset report — node and edge counts (packages, versions, maintainers,
advisories, organizations, repositories, lockfiles;
DEPENDS_ON,AFFECTED_BY,RESOLVES, …). - Query latency — repeats real package, dependents, advisory and full
blast-radius queries N times and reports min / P50 / P95 / P99 / max / mean.
Blast radius is timed end-to-end across every HydraDB query it issues, not
just the
algo.SSPathstraversal. - Precision / recall — for the demo-org applications, compares what the
graph flags (
RESOLVES→AFFECTED_BY) against ground truth from the live Google OSV API, at the(app|package|version)exposure level and the strict advisory-ID level. The exposure level is the fair metric, because the same bug is oftennpm-audit-*in one database andGHSA-*in the other.
Recorded against a locally seeded HydraDB via bun run bench:all --iter 25
(N = 25 real queries per case). Re-run any time to regenerate.
Dataset
| Metric | Count |
|---|---|
| Packages | 56 |
| Package versions | 80 |
| Maintainers | 80 |
| Advisories | 205 |
| Organizations / repositories / lockfiles | 1 / 8 / 8 |
DEPENDS_ON edges |
42 |
AFFECTED_BY edges |
223 |
RESOLVES edges |
50 |
Query latency (ms) — package lookup, dependents and advisory queries are single HydraDB round-trips; blast radius is the full end-to-end cost.
| Query | P50 | P95 | P99 |
|---|---|---|---|
| package lodash@4.17.20 | 4.4 | 5.1 | 6.2 |
| dependents lodash@4.17.20 | 4.5 | 5.7 | 5.9 |
| advisories lodash@4.17.20 | 21.3 | 23.2 | 24.1 |
| blast-radius lodash@4.17.20 (full) | 9.7 | 11.3 | 11.5 |
| blast-radius express@4.18.2 (full) | 4.2 | 5.0 | 5.2 |
| blast-radius request@2.88.2 (full) | 8.9 | 10.1 | 10.4 |
| blast-radius aiohttp@3.8.4 (full) | 3.7 | 5.1 | 5.5 |
Precision / recall vs OSV ground truth (demo-org applications, exposure
level app|package|version)
| Metric | Value |
|---|---|
| Ground-truth exposed app-versions | 5 |
| Predicted exposed app-versions | 5 |
| True positives | 4 |
| Precision | 0.800 |
| Recall | 0.800 |
| F1 | 0.800 |
The single exposure discrepancy (storefront|qs|6.5.2 flagged by the graph
but not OSV; storefront|axios|1.14.1 flagged by OSV but not the graph) is a
real npm-audit vs OSV advisory-coverage difference, not a traversal error.
The brief scores submissions on precision, recall, query latency and cost, with ground truth free from OSV / the GitHub Advisory Database:
- Precision / recall — measured directly against live Google OSV as ground
truth (
bench:all). Reported at both the exposure level and the strict advisory-ID level in the table above. - Query latency — the P50/P95/P99 table above measures real end-to-end
HydraDB round-trips; blast radius is the full traversal cost, not just the
algo.SSPathscall. - Cost — blast radius and dependents are single cheap traversals (single-digit ms on the seeded graph), and ingestion runs once at startup, not on the request path.
The 3-minute submission video walks through, in order:
- The problem — when a package is compromised at 09:00, which of your services are exposed by 09:06? Why that is a graph-traversal question.
- The project — actik models packages, versions, dependencies, advisories, maintainers and the internal apps that consume them as one graph in HydraDB.
- The demo — scan a real repository, view the exposure score and minimal fix set, open an investigation page, traverse the interactive dependency graph, replay a time-travel exposure window, and run a worm-propagation simulation.
- HydraDB — every feature is a graph traversal (
algo.SSPathsblast radius,AFFECTED_BYjoins, temporalRESOLVESedges, reverseDEPENDS_ONclosure). Without the graph there is no transitive exposure answer.
All configuration lives in .env.example. Secrets stay in .env and are never committed. Key variables:
| Variable | Purpose |
|---|---|
FRONTEND_ORIGIN |
Allowed CORS origin (the frontend's public URL) |
HYDRADB_AUTH_TOKEN |
HydraDB auth token (>= 32 chars, required in prod) |
HYDRADB_HTTP_URL |
HydraDB HTTP endpoint (compose default http://hydradb:8443) |
NPM_REGISTRY_URL / PYPI_JSON_URL / OSV_API_URL |
Data source endpoints |
INGESTION_MAX_PACKAGES / INGESTION_MAX_DEPTH / INGESTION_MAX_ADVISORIES |
Ingestion limits |
INGESTION_CONCURRENCY |
Ingestion parallelism |
DEMO_ORG_PATH |
Path to the synthetic organization dataset |
Open source under the MIT License.
Built with HydraDB for Hack Hydra 2026, Track 02A, Repos, Dependencies + Code as Graphs.