Skip to content

perf(shared-fs): bound exact-slot candidate queries and cache - #331

Closed
peerbit-org wants to merge 7 commits into
masterfrom
perf/shared-fs-slot-candidate-cache
Closed

peerbit-org wants to merge 7 commits into
masterfrom
perf/shared-fs-slot-candidate-cache

Conversation

@peerbit-org

@peerbit-org peerbit-org commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

Bound exact-name candidate caching and query the exact naming slot on a cold miss, preserving complete raw histories and the existing naming winner/conflict/tombstone semantics.

  • Replace scattered parent-sweep/reverse-map ownership with one bounded helper.
  • Bound retained slots plus parent markers (4,096), rows (16,384), estimated bytes (8 MiB), and active point fills (64). Negative slots count; oversized histories remain complete but uncached.
  • Keep generation, cache identity/revision, epoch, mutation, and overlay fences. Singleflight callers receive independent shallow arrays.
  • Retain the linear bulk builder and replace repeated same-id relocation filtering with source-slot eviction.
  • Extend the existing resource harness rather than adding a parallel benchmark framework.

Matched 100k evidence

The identical v2 worker ran in four fresh processes against legacy production 5479d691 and bounded production 4529da1c, using the same executed Peerbit 5.3.35 / Documents 15.0.16 / shared-log 16.0.15 cohort. No install or dedupe. These are real disk-backed SQLite candidate-query measurements, not public file writes, mounted throughput, cold disk, cold join, or the held 5.4.0 rebaseline.

Operation Legacy Bounded
Cold hit, 100k distinct names 1,237.49 ms / 100k decoded rows 94.68 ms / 1 decoded row
Cold miss after cache clear 1,034.99 ms / 100k rows 0.571 ms / 0 rows
Three unrelated-mutation rejected fills 917 / 958 / 1,014 ms 0.674 / 0.319 / 0.270 ms
Candidate rows retained after unique first hit 100k 1

The first query includes lazy SQL index/planner creation. Actual prepared SQL/bindings and EXPLAIN confirm composite kind/parentId/name access. Each lookup's exact candidate identity set is checked outside timing; raw reports include source/lock/module hashes and GC/memory checkpoints.

Tradeoffs: A 100k-row single-name naming history exceeds the budget and repeats at roughly 985 ms instead of a cached 0.102 ms. This fixture has 100k distinct node claimants, not 100k content edits to one file. A full sweep after partial warming reads all 100,100 rows (1.46 s for unique), while legacy already holds that parent in memory. New unseen-name creates perform point absence queries (0.156 ms median versus legacy's 0.067 ms after its first full sweep). Full enumeration, oversized-history transient memory, and incoming caller backlog are not bounded by the retained-cache limits. The same-name campaign still sampled 1,181 MiB heap before GC. Do not interpret this as an across-the-board speedup.

See SLOT_CACHE_RESOURCE.md for historical v1 evidence, matched v2 raw fixtures, reproduction, exact limitations, and package sizes.

Validation and hold

  • 18 focused existing/budget/integration tests plus 19 independent adversarial tests pass.
  • Independent backing-map oracle: 100,000 operations across five budgets, 763,672 known slot results and 55,104 complete sweeps checked without mismatch; accounting and retention limits hold.
  • Two v2 harness smoke cases and all four matched 100k runs pass.
  • Library/CLI build, source lint/format, output/package checks pass. Library: 2,241,961 unpacked bytes / 502,708 gzip bytes, within the 2.75 MB budget; test/results artifacts excluded.
  • Full library run: 549 passed / 13 skipped, but one durable-disposal case used embedded retry x1 despite CLI --retry=0. Therefore the strict first-attempt gate failed; the default reporter omitted the original stack.
  • One isolated diagnostic, changing only that case's retry to zero, passed in 10.47 seconds. It does not clear or explain the original failure; no diagnostic patch is included and no timeout was raised.

Keep this PR held. Fresh cross-platform first-attempt validation, the coherent upstream rebaseline, public namespace/mounted workload coverage, and history/enumeration materialization tradeoffs remain gates. No release is authorized by these microbenchmarks alone.

@peerbit-org

Copy link
Copy Markdown
Collaborator Author

Holding this PR without rerunning CI. Exact head 6ce570e003f36606dec0adb886ac469ed89565c6 ran workflow 33932458719 at workflow attempt 1. The cache-focused coverage and all non-portable checks passed, but the persisted lifecycle gate was not clean: Ubuntu persistent-multi-writer failed after Vitest retry x1 (first execution timed out waiting for a replicator; retry lost persistedReceiptPeerSession capability), while Windows passed only after retry x1; macOS passed its first execution. This matches the known upstream receipt/session liveness class and is unrelated to directory-slot indexing, but our acceptance rule requires first-execution success on Ubuntu, macOS, and Windows. No blind rerun and no merge. We will refresh and obtain fresh three-OS evidence after the upstream coherent fix and PR #300.

@peerbit-org

Copy link
Copy Markdown
Collaborator Author

Resource/performance audit: this remains on hold.

Static review found no correctness blocker, and the 10k benchmark does support width-independent warm candidate selection. It does not yet establish production safety for a large namespace.

Before merge, please refresh onto current master and provide one matched, fresh-process 100k unique-name benchmark with:

  • disk-backed data and node --expose-gc;
  • cold hit/miss plus randomized p50/p95/p99;
  • heap/RSS and cache cardinalities before load, after cold fill, after explicit GC, and after eviction;
  • a full-list control;
  • an actual 100k same-name install/query case;
  • mutation-during-cold-fill coverage;
  • forced eviction proving reverse-index cleanup;
  • fresh zero-retry Ubuntu/macOS/Windows CI.

The current design retains a reverse-map entry and placement object per cached naming row. A single huge directory is not bounded by the parent-count limit, and the global slotMutationEpoch may cause unrelated mutations to repeatedly discard an O(width) cold fill. If the measured peak or writer interference is material, consider a bounded/LRU point-query cache or a repairable per-parent generation.

The old-base package delta is modest (+25,512 unpacked bytes, +4,768 compressed), so runtime memory and cold-fill behavior—not package size—are the gate. No CI rerun requested.

@peerbit-org
peerbit-org force-pushed the perf/shared-fs-slot-candidate-cache branch from 6ce570e to 5479d69 Compare September 5, 2026 07:00
@peerbit-org

Copy link
Copy Markdown
Collaborator Author

100k resource campaign and bounded bulk-build fix

Refreshed this branch onto master 403942374a2be0b6c16cfd5467c748cf4900e6ad and added a real SQLite index-backed, fresh-process resource harness. No dependency installation, deduplication, timeout inflation, or CI rerun.

The campaign found an O(N²) same-name bulk constructor: 2,000 rows incurred 4,006,000 identity reads. A 100,000-row same-name campaign exceeded its approximately three-minute budget; that is a censored timeout, not a measured install duration. The fixed bulk constructor takes 119–145 ms in the three observed same-name samples. Existing dynamic upsert/removal semantics remain unchanged.

Independent review compared the actual old/new implementations over 10,000 randomized histories (2,513,641 rows), plus 50,000 randomized install/delete/eviction operations including cross-parent moves. Ordered cache rows, singleton/history representation, reverse placements, and eviction epochs matched.

The remaining limits are material:

  • 100k cold candidate hit/miss still takes roughly 1.05–1.38 seconds; the query materializes the whole parent history.
  • The unique-name cache retains 52.26 MB heap versus 44.72 MB for the former id-map control (+7.54 MB).
  • Three correctly rejected fills after unrelated index writes sampled 1.255 GB heap / 1.544 GB RSS. Explicit GC releases the temporary heap; RSS can remain allocated. This is not the retained cache size or a proven retained-row leak.
  • Moving many existing IDs out of another cached same-name history can still incur quadratic repeated removal. The fix establishes linear bulk grouping, not universal whole-install linearity.

The report and raw JSON fixtures are in packages/shared-fs/SLOT_CACHE_RESOURCE.md. Scope is index-only candidate lookup, not public fs.stat/fs.list, mounted IO, Documents writes, replication, or persisted receipt performance. The executed cohort is Peerbit 5.3.35 / document 15.0.16 / shared-log 16.0.15 / indexer-sqlite3 3.0.19, not the held 5.4.0 refresh.

Local validation: 13 focused cache/bulk/race tests, both fresh-process smoke cases, build/lint/format, package checks, and the two full 100k campaigns passed without retries. Library package remains 83 files / 2,219,037 unpacked bytes; raw fixtures and benchmarks are excluded.

Keep this PR held. Next gate is a bounded exact-slot query/cache and query-materialization policy that preserves same-ID replacement safety across names and parents, followed by this same campaign and filesystem-level validation on the next upstream cohort. This update does not unblock release #324.

@peerbit-org

peerbit-org commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Exact head 5479d69148206cf60e0d2ff297cde8abcc614e0e now passes the strict first-attempt CI gate. All five naturally triggered workflows succeeded, with no manual rerun.

Ubuntu, macOS and Windows each passed 525 library tests and 33 CLI tests. Durable disposal passed 22/22 and persistent multiwriter 1/1 on the first attempt on every OS. Raw logs have no retry markers, failed tests, receipt-session failures or replicator timeouts. Native, interop, package installs, formatting, changesets and deployment validation are also green.

Shared FS CI

Still held: this verifies cross-OS correctness checks on this branch's locked Peerbit 5.3.35 cohort, not the 100k cold-query/memory gate or the held 5.4.0 confirmation/session issue. Keep the bounded exact-slot/materialization work and the existing release hold; no merge or release performed.

@peerbit-org peerbit-org changed the title perf(shared-fs): index warm directory slots by exact name perf(shared-fs): bound exact-slot candidate queries and cache Sep 5, 2026
@peerbit-org

Copy link
Copy Markdown
Collaborator Author

Additional public-API validation at exact head 9cde97620e0966b120a050631c1a9cf014cb57b9: the existing opt-in slot-candidate-cache.bench.test.ts passed once with --retry=0 (14.24-second test, 15.26 seconds overall). No source changes, timeout inflation, dependency install or dedupe.

Real openSharedFs + mkdir + writeBatch fixtures, followed by candidate-cache clear and 100 repeated two-segment fs.stat calls:

Directory width Stat p50 Stat p95 Candidate rows examined across 100 stats Retained slot histories / reverse entries
100 0.0165 ms 0.0357 ms 200 2 / 2
10,000 0.0107 ms 0.0193 ms 200 2 / 2

The structural result is the useful gate: directory width did not increase examined or retained candidate rows for this path. Small-sample timing differences are descriptive, not evidence that wider directories are faster. This workload uses one in-process peer, the existing 5.3.35 cohort, and no mount or remote persisted receipts. It does not clear the full-suite retry, oversized-history/full-list, or latest-cohort holds.

@peerbit-org

peerbit-org commented Sep 5, 2026 •

Copy link
Copy Markdown
Collaborator Author

Strict CI audit for exact head 9cde97620e0966b120a050631c1a9cf014cb57b9: keep held.

Shared FS CI run 33954049406 completed successfully on workflow attempt 1, but workflow success is not a first-attempt test pass.

OS Library / CLI Raw-log result
Ubuntu 549 passed, 13 skipped / 33 passed No test retry or failed-test marker
Windows 549 passed, 13 skipped / 33 passed No test retry or failed-test marker
macOS 549 passed, 13 skipped / 33 passed Persistent-multi-writer used retry x1

The macOS case was keeps three authenticated disk replicas converged through writes, restart, conflict, and disposal, reported at 154,628 ms in job 101274007373. The default reporter retained eventual metrics and the retry marker, but no first-attempt failure stack. Do not attribute the cause to cache or upstream from this evidence.

The naming-conflict/tombstone disposal case passed without retry in each CI job (Ubuntu 14.705 s, macOS 10.289 s, Windows 11.735 s); its separate earlier local full-suite retry remains recorded. No workflow rerun, timeout inflation, or retry-until-green was requested.

All resource/history/latest-cohort gates in the PR body remain. A useful separate downstream cleanup is first-failure reporting for embedded retry-enabled tests, so future failures retain actionable phase/stack evidence.

Final overall checks: Shared FS CI, cross-OS interop, Changesets and all three tarball installs passed. Hosting validation failed during dependency installation, before a shared-fs build/test result: node-datachannel@0.32.3 / prebuild 13.0.1 warned This package does not support N-API version undefined, then threw TypeError: expected first argument to be an array from each-series-async/index.js:21:8 / prebuild/bin.js:48:3 and exited 2. This is a separate unresolved installation failure, not a passing check or an established cache regression.

@peerbit-org

Copy link
Copy Markdown
Collaborator Author

Downstream follow-through; this PR remains held at exact head 9cde976.

  • Browser-only install cleanup #341 merged after fresh first-attempt green hosting validation (530 tests, 7 browser startups, 7 Worker dry-runs). It does not retroactively clear this head's old install failure.
  • #340 now preserves first-attempt test stacks and rejects actual retries. Its first Ubuntu run exposed a missing persisted-receipt session in the restarted-writer readiness preflight; strict CI remains red. No rerun/check bypass or release.
  • An isolated public-API namespace experiment is published at 66d42a48, stacked on this head, not proposed for merge. Six fresh-process local scale cases passed. Warm point stat stayed ~0.013ms at 100 and 10k files, but full 10k listing's three-call median was ~995ms. Repeated content saves and rename/delete/recreate history have distinct costs; full methodology/counters/cohort hashes and limitations are in the report.
  • The real three-peer partition/rejoin case converged on claims/winner/bytes, then failed the 90-second child shutdown cap despite fulfilled stop promises. A separate native sample reaches Node RunCleanup -> CleanupHandles -> uv_run after beforeExit. Upstream handoff is prepared; no forced-exit workaround or claim of latest-cohort reproduction.

The next small cleanup candidate is unused recursive DAG depth computation: no consumer reads it, while winner ordering uses stored causalDepth. It needs head/winner invariance tests and measurements, not a larger cache. No new production optimization is part of the experiment branch.

@peerbit-org

Copy link
Copy Markdown
Collaborator Author

Review against the new master (Peerbit 5.4.6 base, #352/#353): hold, changes needed

This merges cleanly and I found no correctness bug in cache coherence. Renames, deletes, same-id moves, GC removals, overlay reads and reopen are all fenced by the generation, identity, revision, slot:P and mutation-epoch checks. Three independently verified problems block landing as-is:

  1. installSweep evicts before it checks capacity. evictParent(parentId) and relocate run before fits(1 + prepared.length, rows, bytes). For a directory with ≥4,096 distinct names (or about 9k naming rows by the byte estimate), every list() throws away that directory's still-valid exact-slot lookups and then refuses to cache it. Alternating list/stat on such a directory never hits the cache. Fix: compute names, rows and bytes first, and return false without touching existing slots when it cannot fit. Add a regression test for list → stat → list → stat with no second row query.

  2. Complete-directory listings share one 4,096-entry / 8 MiB LRU with point lookups. master keeps an unbounded per-directory sweep cache. With this PR:

    • watch rebuilds that walk a subtree larger than the cache (fullDiff/listByParentId after any spine-parent event) miss on every flush;
    • a file watch in a directory with ≥4,096 names re-lists it uncached on every sibling event;
    • two 2,500-name directories listed alternately evict each other.

    Either give complete listings their own scan-resistant, byte-based budget, or measure watch-rebuild and mount-readdir workloads on 5.4.6 and state the tradeoff.

  3. Semantic conflict with perf(shared-fs): bound cache epoch metadata #303. perf(shared-fs): bound cache epoch metadata #303's cache-epoch-bound.test.ts reads program.slotSweepCache (three assertions), which this PR deletes. Once both land, in either order, that test throws TypeError. Plan: land perf(shared-fs): bound cache epoch metadata #303 first. Then, when rebasing this PR, keep its open() block (which already has the monotonic cacheGlobalEpoch increment), port the three assertions to slotCandidateCache (for example getSweep(parentId) === undefined), and update perf(shared-fs): bound cache epoch metadata #303's README sentence to name the separate slot-cache limits.

Also:

  • the six fixtures/slot-resource-*.json files (2,143 lines) are not read by any test;
  • SLOT_CACHE_RESOURCE.md reports runs on the replaced 5.3.35 cohort.

Dated reports now live on the evidence branch, not the package root. After those changes, this needs fresh three-OS first-attempt CI on the 5.4.6 base; the last evidence at 9cde9762 was on 5.3.35 and needed a retry on macOS.

@peerbit-org

Copy link
Copy Markdown
Collaborator Author

Superseded by #359. It keeps this PR's exact-slot queries and bounded point cache and addresses the review above: it leaves the per-directory listing cache unchanged, checks capacity before evicting, keeps #303's test working, and drops the fixtures. It also adds FIFO admission and a width gate, so filesystems without wide directories never create the extra SQLite indexes (the owner's decision). The branch is kept.

peerbit-org added a commit that referenced this pull request Sep 27, 2026
Path resolution, stat and create checks used to read the whole parent
directory (a sweep) to find one name. In a wide directory that has not
been listed, a cold lookup therefore cost O(width).

Resolve a single (parentId, name) slot in this order:
1. a cached directory sweep (slotSweepCache, unchanged from master),
   which also proves absence with zero queries;
2. a separate bounded point cache of exact slot histories, including
   cached negatives (at most 4,096 slots and parent markers, 16,384 rows,
   about 8 MiB estimated, 64 distinct in-flight queries);
3. one exact kind+parentId+name index query, installed with the same
   fences the sweep fill uses: open generation, cache identity and the
   per-directory slot:<parentId> epoch. No global mutation epoch.

The point cache decides admission before it mutates anything: a history
too large to retain is returned whole, not cached, and evicts nothing.
Change events keep cached slots exact, relocating same-id replacements
by evicting the source slot. Directory listings and slotSweepCache are
not changed. open(), retireOverlay() and close() replace both caches.

Based on the point tier of #331 without its complete-parent tracking.
peerbit-org added a commit that referenced this pull request Sep 27, 2026
Env-gated (PEERBIT_SHARED_FS_SLOT_CACHE_BENCH=1); skipped in normal runs. Adapted from #331's slot-candidate bench: an end-to-end case over writeBatch-built directories and a seeded 100k-row SQLite index case. It only uses fields present on master too, so the same file measures a baseline. Gates are structural (query kinds and counts); timings are descriptive.
peerbit-org added a commit that referenced this pull request Sep 27, 2026
…359)

Lookups in unlisted directories read them with a bounded query (at most
2,049 rows, same index as the listing sweep) and cache narrow ones whole
as before. Directories found or listed wider than 2,048 naming rows use
exact (parentId, name) queries with a bounded point cache. Filesystems
without wide directories never create the slot indexes. Point-tier
waiters are admitted per key in FIFO order; sweep fills are fenced by
cache identity. Cold hit at 100k rows: 1,147 ms -> 337 ms; later cold
miss 913 ms -> 0.23 ms; bulk creates at parity. Supersedes #331.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant