From 95e68b7a41bd380619b2353eb8456b71a814092d Mon Sep 17 00:00:00 2001 From: joyful-ii-V-I Date: Mon, 21 Sep 2026 14:55:36 -0400 Subject: [PATCH 1/3] docs(prompts): every COBOL number we hold came from 716 public files, and nobody had written the round that measures more The COBOL stack behind issue #70 is measured entirely on 716 public files -- NIST, CardDemo, an OMP course set, and vendor and exercism samples. That is a convenience sample, not a census, and the roadmap items it cannot see (CICS LINK/XCTL, SQL INCLUDE, COPY REPLACING, nested programs, free format) are currently ordered by guesswork. prompts/cobol-measure-on-your-corpus.md is the round that fixes that, runnable by anyone with public COBOL to point it at. It carries the facts a reader needs to reproduce the recommended configuration -- the two-pass two-grammar merge, the two scanner bugs pass A needs patched, the token-splitting bug in pass B, the role merge, and the best-tree trap that looks better while collapsing PERFORM recall to 54.5% -- and asks for parse rates by artifact kind and recall SPLIT BY TIER, because tree edges and token-fallback edges are different products at identical totals. A separate, optional section is for a reader with production COBOL they cannot share. It is aggregate-only by construction: counts, rates and histogram buckets, never a file name, a symbol, a snippet or a path, and any step that cannot be answered without naming something is skipped and the skip reported. Part 2 is the decision the fork turns on: whether their extraction already knows where a symbol's BODY ends, which is what serving fetch_body honestly from a SCIP index requires (typed_enclosing_range, single_line 10 / multi_line 11; enclosing_range 7 is deprecated). Either answer is useful and neither is a heuristic -- we will not guess where a symbol ends. prompts/README.md gains its row and README.md's prompt count moves to thirteen, which readmedriftcheck (I1) re-derives from the directory. Gates: readmedriftcheck ALL PASS, deckcheck ALL PASS, ripwirepubliccheck ALL PASS. Co-Authored-By: Claude Opus 5 --- README.md | 4 +- prompts/README.md | 1 + prompts/cobol-measure-on-your-corpus.md | 259 ++++++++++++++++++++++++ 3 files changed, 262 insertions(+), 2 deletions(-) create mode 100644 prompts/cobol-measure-on-your-corpus.md diff --git a/README.md b/README.md index fdb1e5801..1790391ae 100644 --- a/README.md +++ b/README.md @@ -1990,7 +1990,7 @@ it points at the installer's staged copy of the skills when the cwd is not a che ## Improve it with your agent -[`prompts/`](prompts/) holds twelve **self-contained orchestrator prompts**: the loops this project is +[`prompts/`](prompts/) holds thirteen **self-contained orchestrator prompts**: the loops this project is built with, written so a coding agent can run them. They encode the workflow rather than describing it. @@ -2787,7 +2787,7 @@ files under `bench/`. ### 16. Improvement -`prompts/` holds twelve **self-contained orchestrator prompts**. Each prompt is a workflow that a +`prompts/` holds thirteen **self-contained orchestrator prompts**. Each prompt is a workflow that a coding agent can run against this repository. Each prompt writes a plan and stops for your approval before it runs a command. Build the binary first. The prompts measure against the binary. diff --git a/prompts/README.md b/prompts/README.md index aa58731a2..cff2613be 100644 --- a/prompts/README.md +++ b/prompts/README.md @@ -10,6 +10,7 @@ before you approve it. Read the plan, cut what you disagree with, then say go. | Prompt | Who it is for | What it produces | | --- | --- | --- | | [`add-a-language.md`](add-a-language.md) | Anyone who wants a language ripwire does not index yet | A vendored grammar, extraction, disclosed blind spots and a red-first gate — the path the Elixir grammar actually took, file by file, including the parse-rate measurement that decides whether to start at all | +| [`cobol-measure-on-your-corpus.md`](cobol-measure-on-your-corpus.md) | Anyone with COBOL to point it at — public IBM i corpora first, and optionally production code they cannot share, which is a separate, aggregate-only section | Parse rates by artifact kind and PERFORM/CALL/COPY recall **split by tier**, tree edges apart from token-fallback edges, for the two-pass two-grammar stack; occurrence counts for the constructs our roadmap is still guessing at (CICS, embedded SQL, COPY REPLACING, nested programs, free format); and a yes/no on whether a SCIP resolver can produce `typed_enclosing_range`, which is what decides whether the fork can go away (#70). | | [`improve-for-my-language.md`](improve-for-my-language.md) | Anyone running ripwire on their own repository — a first session included; the maintainer plan is a labelled section at the end | Transcript-grounded gaps for one language — grammar coverage, symbol kinds, ranking, legends, first-run friction — as an issue-ready report with a reproduction per finding. | | [`improve-quality-panel.md`](improve-quality-panel.md) | Anyone whose panel shortlist they can judge — run it from your OWN codebase | Per-family agreement verdicts from blind reads of real functions, misses diagnosed by pipeline stage, and an ordered plan in which any ranking change owes a pre-registered calibration round. | | [`full-audit.md`](full-audit.md) | A maintainer, anyone deciding whether to trust the tool, or anyone willing to run it against their own large repository | A severity-ranked, gated audit across six lenses — bugs and hostile inputs, measured performance at the scale rung, verb-to-moment matching, token efficiency, an ecosystem scan of papers and repos with real momentum, and the honesty of the output — after first proving each instrument can see what it measures. | diff --git a/prompts/cobol-measure-on-your-corpus.md b/prompts/cobol-measure-on-your-corpus.md new file mode 100644 index 000000000..03ab8919d --- /dev/null +++ b/prompts/cobol-measure-on-your-corpus.md @@ -0,0 +1,259 @@ +# Measure the COBOL stack, then decide what to build + +This is the COBOL round of [issue #70](https://github.com/redhat-et/ripwire/issues/70), written so +that anyone can run it. **Most of it needs nothing but public code**: the two-pass COBOL stack +described below, pointed at public IBM i corpora, producing parse rates, recall split by tier, and +frequency counts for the constructs our roadmap is currently guessing at. That is the part that was +volunteered for, and it is the spine of this prompt. + +One section — clearly marked, entirely optional — is for a reader who also has production COBOL they +cannot share. **That section is aggregate-only, and the constraint is absolute: counts, rates and +histogram buckets leave your machine, nothing else.** No file names, no program or paragraph names, +no source lines, no paths, no client identifiers, not once. We cannot see your code, we do not want +to see it, and no step here is worth a single identifier. If a step cannot be answered without +naming something, **skip it and report the skip** — a reported skip is a result, and a leak is not. + +Everything measured on the public 716 files is quoted below so you can reproduce it, disagree with +it, or find the case that breaks it. **The measurement comes before any grammar or ingestion code.** +That is the same rule [`add-a-language.md`](add-a-language.md) opens with, and it is not ceremony: +the parse rate and the tier split are what decide whether this work is worth starting and in which +order. + +--- + +## Setup — where to start from + +Base the work on **`main`**, which is **v0.6.2** (tagged 2026-09-21, commit `15a20855`). There is no +branch to wait for; `main` is the base, and a fork of it is the right shape for this. + +```bash +cmake -S . -B build && cmake --build build -j +python3 test/pargates.py . ./build/ripwire -j 6 # the gate suite, in the foreground +``` + +Do not add a build type. `-DCMAKE_BUILD_TYPE=Release` defines `NDEBUG`, which compiles the +degrade-path diagnostics out and blinds the gates that assert them. + +Record `ripwire --version` and the `` line from `ripwire --doctor` in your report +header, so every number below has a binary attached to it. + +--- + +## What has already been measured, and on what + +Every COBOL number this project holds came from **716 public files, 27.8 MB**: NIST, CardDemo, an +OMP course set, and IBM, Microsoft and exercism samples. That is a convenience sample, not a census, +and its shape is the reason this round exists — sample code under-represents the constructs real +shops live in. + +The recommended configuration is a **two-pass, two-grammar merge**: + +- **Pass A — `yutaro-sakamoto/tree-sitter-cobol`, the PR #41 branch, patched.** It carries two + scanner bugs you will hit immediately: an out-of-bounds read in `start_with_word` that a sanitizer + build catches, and an infinite loop at end of file when the last line is under six columns — a + one-byte file hangs. Patch both before you measure anything, and say in your report that you did. +- **Pass B — `barrettotte/treesitter-ibmi`** (MIT; COBOL, DDS, RPG and CL in one grammar). Its + compiled object is 23 KB against pass A's 16.4 MB. Known bug: a token splits words, producing + `P` + `ERFORM`; the workaround is to glue adjacent tokens before matching. +- **Merged BY ROLE, not by best tree.** Structure from one pass, edges from the other. +- **Plus tier-2 edges**: PERFORM, CALL and COPY targets read from pass B's error-free *token* + stream, used only for paragraphs whose pass A tree failed. Two false edges over the 716 files. +- **Plus an offset-preserving reference-format normaliser** — it blanks the sequence and indicator + areas without moving a single byte offset, so every span the extraction reports still points at + the real file. + +**The best-tree-per-file merge is a trap, and it is the most important thing on this page.** Picking +whichever pass produced the cleaner tree for each file *looks* better — 95% of files come out clean +— and it destroys the answer: PERFORM recall collapses to **54.5%**, because barrettotte trees carry +no statement nodes at all. A file can be clean and edge-free at the same time. If your own merge +strategy is scored on tree cleanliness, you are measuring the wrong thing. + +Measured on the public 716, strict-clean, against a hand-checked set of 182 untuned rows: + +| Quantity | Measured | +| --- | --- | +| Definitions | 98.7 / 100 | +| PERFORM edges | 100 / 100 | +| CALL edges | 100 / 100 | +| COPY edges | 100 / 100 | +| Files fully covered | 96.7% | +| Programs parsed | 98.9% | +| Copybooks parsed | 98.9% | +| Cost | 1.76 s, against 1.03 s for the single pass, over 27.8 MB | + +The stack emits a per-file disclosure record, and **the aggregate of that record is the report this +prompt is asking for**: whether pass A was clean, including zero-width hidden errors; which +definitions came from pass B; which paragraphs have unknown edges; a TIER attribute on every edge +(tree or token); bytes neither pass parsed; dynamic CALL sites left unresolved; and what the +normaliser blanked. + +--- + +## Part 1 — the stack against public COBOL + +Run the configuration above over as much public IBM i COBOL as you can assemble, and report the +numbers below. **Extending the public corpus is itself a result**: our 716 files are what one person +could find, and a second person's list is the only way to learn what that search missed. Name the +corpora you added and where they came from, so the next run can reproduce yours. + +### 1.1 — establish ground truth cheaply + +You cannot report recall without a truth set, and you do not need an expensive one. + +This project built its truth with a **lexical oracle**: a small script that scans the raw bytes for +PERFORM, CALL and COPY and records every target it can see, independent of any parse tree. That +produced 105,906 rows, and the rows were hand-checked. Do the same thing, and then stop: **a +hand-checked sample of a few hundred rows is enough to place a recall number**, and building more +instrument than that is the classic way this round turns into a project. Sample across artifact +kinds rather than taking the first few hundred rows of one file. + +Say how many rows you checked and how you sampled them. A recall figure without its denominator and +its sampling rule is an impression. + +### 1.2 — the numbers + +Report each of these, split as indicated. Every one is a count or a rate. + +1. **Clean-parse rate, split by artifact kind** — programs, copybooks, DDS. One rate per kind, with + its denominator. A single blended rate hides the kind that is actually broken. +2. **Definitions found**, and the **false-definition rate** — a definition the stack reports that + your oracle says is not one. The false rate matters more than the found rate: a map that invents + symbols is worse than one that misses them. +3. **PERFORM, CALL and COPY recall, SPLIT BY TIER.** Tree-tier edges and token-tier edges reported + separately, never summed into one number. **This split is the point of the whole exercise** — it + says how much of the answer rests on the token fallback, and a stack whose recall is carried by + tier-2 is a different product from one whose recall is carried by trees, even at identical totals. +4. **Bytes neither pass parsed.** Absolute and as a fraction of corpus bytes. +5. **Dynamic CALL count** — call sites whose target is an identifier rather than a literal, left + unresolved. This is a floor, not a total; label it as one. +6. **Normaliser blanks** — how many files had sequence or indicator content blanked, and how many + bytes. A normaliser that blanks nothing on a corpus is not being exercised. + +### 1.3 — the frequency question, which is the highest-value thing here + +Our gap list for COBOL is **guesswork about what real IBM i shops use**. It is: + +- `EXEC CICS LINK` / `XCTL` +- SQL `INCLUDE` +- `COPY … REPLACING` +- nested programs +- free format + +**Report how often each one appears** — a count of occurrences and a count of files containing at +least one, per corpus. That reorders our roadmap directly, and no amount of reading the standard +substitutes for it. + +Run this on the public corpus and expect the answer to be skewed: **public sample code +under-represents CICS and embedded SQL badly**, because teaching material and demo applications are +written to run without a transaction monitor or a database. Saying so with numbers is more useful +than the numbers alone — a near-zero CICS count on public code is evidence about the corpus, not +about COBOL. + +--- + +## Part 1b — optional: a corpus you cannot share + +Skip this section entirely unless you have access to production COBOL under an agreement that +forbids sharing it. Nothing in Part 1 depends on it. + +If you do have that access, here is the honest framing: **real client code is the only population +that tests this stack, and nobody outside your client base can measure it.** Our 716 public files +are teaching material and demos. Whatever breaks on forty years of accreted production code is +invisible to us and always will be. + +**The rule for this section, repeated because it is the only thing that makes it runnable:** + +- **Aggregate counts only.** Every figure is a count, a rate, or a histogram bucket. +- **Never** a file name, a program name, a paragraph name, a copybook name, a symbol, a source line, + a path fragment, a library or schema name, or anything identifying a client. +- Histogram buckets, not distributions with outliers attached — "11 files in the 10–50 KB bucket", + never "the largest file". +- **If a step cannot be answered without naming something, skip it and report the skip.** Write the + skip down: "step 1.2(5) skipped, cannot be answered in aggregate on this corpus" is a finding we + can act on. An identifier is not. +- Run the numbers, read the output yourself, and send only the table. Do not paste tool output. + +Report the same list as 1.2, plus the same frequency counts from 1.3 — **the frequency question +belongs in both sections**, and the two answers being different is exactly what we want to see. A +public corpus that says CICS is rare and a production corpus that says it is everywhere is a single +finding worth more than either number alone. + +--- + +## Part 2 — can your resolver produce `typed_enclosing_range`? + +This is a design question, not a measurement, and it is the **single blocker** for issue #70's actual +ask. Answer it before writing extraction code, because the answer decides which extraction code is +worth writing. + +The situation, precisely. In SCIP, a definition occurrence's `range` is the **identifier**, not the +body — it covers the name and stops. Serving `fetch_body` from a SCIP index therefore needs a second +range, the one that spans the whole symbol. SCIP spells that `typed_enclosing_range`: + + single_line_enclosing_range = 10 + multi_line_enclosing_range = 11 + enclosing_range = 7 // deprecated, do not emit + +So the question for your extraction is narrow and answerable: **does it already know where a +symbol's body ends?** + +- **If it does** — if the point at which a program, paragraph or section ends falls out of your parse + or your resolution logic — then the SCIP-source path is viable. ripwire consumes the index, serves + bodies honestly from the enclosing range, and **the fork can go away**. That is the outcome + everyone wants, and it is contingent on this one fact. +- **If it does not** — if end-of-body is something your pipeline never computes — then a first-class + vendored grammar is the only honest route, and we should stop designing a SCIP-source path on your + behalf. That is not a worse answer; it is a cheaper one, arrived at before either side builds + against an assumption. + +**We will not ship a heuristic that guesses where a symbol ends.** A body served from a guessed +range is wrong in a way the reader cannot detect, which is the one failure mode this project treats +as disqualifying. So "we could approximate it" is a *no* for the purposes of this question — answer +for what your extraction knows, not what it could be made to infer. + +### 2.1 — re-run your verb telemetry + +Previously reported: **17 sessions, 317 MCP calls, eight native verbs** — `fetch_body`, `grep`, +`find_symbol`, `find_referencing_symbols`, `uses`, `explore`, `for`, `analyze` — and **none** of the +parse-tree features. That mix is why `typed_enclosing_range` is the blocker rather than one item +among many: `fetch_body` is in the top of it. + +Re-run the same count on the current release and report the same two numbers plus the verb +histogram. If the mix moved, that changes which surface is worth hardening; if it did not, eight +verbs is a much smaller contract than the fork is currently carrying, and that is an argument for +retiring the fork sooner rather than later. + +--- + +## The report + +One document, ordered: setup header, Part 1 numbers, Part 1b numbers if you ran it, the Part 2 +answer. Then: + +- **Every number carries its denominator and its instrument.** "98.7 / 100 on hand-checked rows + sampled across three corpora" is a result; "definitions are good" is not. +- **A zero is a measurement; absent is not zero.** A count that cannot be a total is a floor and + says so. Dynamic CALL is a floor. Bytes-not-parsed is a total. +- **Report the losses in their own section**, not folded into an average. The corpus where the parse + rate fell is the interesting one. +- **Name every skip.** Especially in Part 1b, where a skip is the correct answer to several steps. +- Post the Part 1 and Part 2 results on + [issue #70](https://github.com/redhat-et/ripwire/issues/70). Part 1b's aggregate table goes in the + same place — it contains no code and never will. + +### Honesty rules + +- **Never publish a number without an instrument that pins it.** "Works well on real COBOL" is not a + result; a rate against a hand-checked sample with a stated sampling rule is. +- **A merge strategy is judged on recall, not on tree cleanliness.** The best-tree trap above is the + worked example; assume your own variant has a version of it and go looking. +- **Tier is never summed away.** Tree edges and token edges are reported apart, in every table where + either appears. +- **A disclosure surface that under-reports is worse than one that is absent.** If your stack cannot + tell whether a paragraph's edges are known, say that, rather than emitting a confident zero. + +--- + +**Write the plan, then STOP for my go-ahead.** Say which corpora you will measure, whether you are +running Part 1b at all, and which of the two Part 2 answers you already suspect — that last one is +cheap to state and expensive to discover late. From 787fe55aa3016db4c7eefc44dc057a0c767a2c0e Mon Sep 17 00:00:00 2001 From: joyful-ii-V-I Date: Mon, 21 Sep 2026 14:58:03 -0400 Subject: [PATCH 2/3] docs(readme): the prompt list said seven and named seven of ten MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Three prompts were missing from the collapsed list — add-a-language and improve-quality-panel were already absent at twelve, and this lane's COBOL round makes three. Nothing gates that sentence, which is why it drifted; the gated count above it was right the whole time. Now: three highlighted plus ten others equals the thirteen files in prompts/. Co-Authored-By: Claude Opus 5 --- README.md | 8 +++++--- 1 file changed, 5 insertions(+), 3 deletions(-) diff --git a/README.md b/README.md index 1790391ae..6b59f1e1c 100644 --- a/README.md +++ b/README.md @@ -2018,11 +2018,13 @@ Three worth starting with: | [`capture-audit.md`](prompts/capture-audit.md) | A fresh showcase capture read by parallel adversarial lenses, and the findings turned into family-wide gates. |
-The other seven — head-to-head, ranking-eval from your own sessions, per-language, onboarding, sibling sweep, command tour, showcase build +The other ten — head-to-head, ranking-eval from your own sessions, per-language, onboarding, sibling sweep, command tour, showcase build, add a language, quality-panel calibration, COBOL corpus measurement -The other seven — a paired head-to-head against a competitor, a ranking-eval loop that mines real +The other ten — a paired head-to-head against a competitor, a ranking-eval loop that mines real retrieval misses from your own sessions, a per-language improvement pass, a zero-context onboarding -study, a sibling sweep, a live command tour, a showcase build — are listed with their audiences in +study, a sibling sweep, a live command tour, a showcase build, the path a new language's grammar +actually took, a quality-panel calibration round, and a COBOL two-pass corpus measurement — are +listed with their audiences in [`prompts/README.md`](prompts/README.md). Each states its own scope and its honesty rules, and most name the gates they must leave green.
From b4ff2a047290c74283fc829d143699f842d9a840 Mon Sep 17 00:00:00 2001 From: joyful-ii-V-I Date: Tue, 22 Sep 2026 06:43:58 -0400 Subject: [PATCH 3/3] fix(prompts): strip unsourced COBOL numbers, pin commits, define the frequency count MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit review of #319: several figures in the "already measured" table and prose (27.8 MB corpus size, 23 KB/16.4 MB compiled-object sizes, 98.7/100 definitions, 98.9% programs/copybooks parsed, 1.76s/1.03s cost, the 105,906-row lexical-oracle count, the 182-row hand-checked sample) do not appear anywhere in issue #70's public thread or the PR description. Removed them; kept only what the thread actually supports (716 files, 96.7% files fully covered, 100% combined PERFORM/CALL/COPY recall and precision) and said plainly that the finer splits are what Part 1 re-derives, denominators included. Also, from CodeRabbit's review: - require pinning the exact commit for both grammar passes and including the patch content, so "patched" is reproducible (not just a moving branch) - add a minimal definitions oracle (PROGRAM-ID/paragraph-name/SECTION scan) so 1.2(2)'s false-definition rate has a defined procedure - define the frequency-counting rule for 1.3 (raw-text literal match, comment lines and string literals skipped, per-occurrence vs per-file counted) - README §16's collapsed prompt-category list still named only the pre-#319 ten; added the three it was missing (add-a-language, improve-quality-panel, cobol-measure-on-your-corpus) so it doesn't undercount against "thirteen" Gates run from this worktree: ripwirepubliccheck.sh ALL PASS (2771 files swept), readmedriftcheck.sh ALL PASS (arm I1 still reads thirteen prompts against the 13 files). Co-Authored-By: Claude Sonnet 5 --- README.md | 3 +- prompts/cobol-measure-on-your-corpus.md | 61 +++++++++++++++---------- 2 files changed, 40 insertions(+), 24 deletions(-) diff --git a/README.md b/README.md index 6b59f1e1c..fa77dc16d 100644 --- a/README.md +++ b/README.md @@ -2795,7 +2795,8 @@ before it runs a command. Build the binary first. The prompts measure against th The prompts cover a full audit, a gap analysis from real use, a capture audit, a head-to-head comparison, a ranking evaluation, a language improvement pass, a zero-context onboarding study, a -sibling sweep, a command tour, and a presentation build. The index is `prompts/README.md`. +sibling sweep, a command tour, a presentation build, adding a new language, a quality-panel +calibration round, and a COBOL corpus measurement. The index is `prompts/README.md`. If the tool answers incorrectly on your codebase, run `prompts/improve-for-my-language.md`. The prompt harvests your session transcript and produces one finding per event with its evidence. Open diff --git a/prompts/cobol-measure-on-your-corpus.md b/prompts/cobol-measure-on-your-corpus.md index 03ab8919d..8d67a4b31 100644 --- a/prompts/cobol-measure-on-your-corpus.md +++ b/prompts/cobol-measure-on-your-corpus.md @@ -41,7 +41,7 @@ header, so every number below has a binary attached to it. ## What has already been measured, and on what -Every COBOL number this project holds came from **716 public files, 27.8 MB**: NIST, CardDemo, an +Every COBOL number this project holds came from **716 public files**: NIST, CardDemo, an OMP course set, and IBM, Microsoft and exercism samples. That is a convenience sample, not a census, and its shape is the reason this round exists — sample code under-represents the constructs real shops live in. @@ -51,13 +51,15 @@ The recommended configuration is a **two-pass, two-grammar merge**: - **Pass A — `yutaro-sakamoto/tree-sitter-cobol`, the PR #41 branch, patched.** It carries two scanner bugs you will hit immediately: an out-of-bounds read in `start_with_word` that a sanitizer build catches, and an infinite loop at end of file when the last line is under six columns — a - one-byte file hangs. Patch both before you measure anything, and say in your report that you did. -- **Pass B — `barrettotte/treesitter-ibmi`** (MIT; COBOL, DDS, RPG and CL in one grammar). Its - compiled object is 23 KB against pass A's 16.4 MB. Known bug: a token splits words, producing - `P` + `ERFORM`; the workaround is to glue adjacent tokens before matching. + one-byte file hangs. Patch both before you measure anything. **Record the exact commit you forked + from, and include or link the patch content itself, in your report** — a branch moves, and + "patched" without a pinned commit and a diff is not reproducible. +- **Pass B — `barrettotte/treesitter-ibmi`** (MIT; COBOL, DDS, RPG and CL in one grammar). Known + bug: its `picture` token splits words such as `PERFORM`; the workaround is to glue adjacent tokens + before matching. Record its commit too — the same reproducibility requirement as pass A. - **Merged BY ROLE, not by best tree.** Structure from one pass, edges from the other. - **Plus tier-2 edges**: PERFORM, CALL and COPY targets read from pass B's error-free *token* - stream, used only for paragraphs whose pass A tree failed. Two false edges over the 716 files. + stream, used only for paragraphs whose pass A tree failed. - **Plus an offset-preserving reference-format normaliser** — it blanks the sequence and indicator areas without moving a single byte offset, so every span the extraction reports still points at the real file. @@ -68,18 +70,17 @@ whichever pass produced the cleaner tree for each file *looks* better — 95% of no statement nodes at all. A file can be clean and edge-free at the same time. If your own merge strategy is scored on tree cleanliness, you are measuring the wrong thing. -Measured on the public 716, strict-clean, against a hand-checked set of 182 untuned rows: +Measured on the public 716, strict-clean, against a hand-checked lexical reference: | Quantity | Measured | | --- | --- | -| Definitions | 98.7 / 100 | -| PERFORM edges | 100 / 100 | -| CALL edges | 100 / 100 | -| COPY edges | 100 / 100 | +| PERFORM, CALL and COPY edges | 100% recall and precision (combined, against the lexical reference) | | Files fully covered | 96.7% | -| Programs parsed | 98.9% | -| Copybooks parsed | 98.9% | -| Cost | 1.76 s, against 1.03 s for the single pass, over 27.8 MB | + +The definition accuracy, the parse rate split by artifact kind, and the wall-clock cost were not +carried forward from the original round with an instrument attached to them, so this page does not +restate them as numbers here. That is exactly what Part 1 below re-derives, with a denominator on +each one — including the hand-checked sample size, which was not preserved either. The stack emits a per-file disclosure record, and **the aggregate of that record is the report this prompt is asking for**: whether pass A was clean, including zero-width hidden errors; which @@ -101,14 +102,20 @@ corpora you added and where they came from, so the next run can reproduce yours. You cannot report recall without a truth set, and you do not need an expensive one. This project built its truth with a **lexical oracle**: a small script that scans the raw bytes for -PERFORM, CALL and COPY and records every target it can see, independent of any parse tree. That -produced 105,906 rows, and the rows were hand-checked. Do the same thing, and then stop: **a -hand-checked sample of a few hundred rows is enough to place a recall number**, and building more -instrument than that is the classic way this round turns into a project. Sample across artifact -kinds rather than taking the first few hundred rows of one file. +PERFORM, CALL and COPY and records every target it can see, independent of any parse tree. The rows +were hand-checked. Do the same thing, and then stop: **a hand-checked sample of a few hundred rows is +enough to place a recall number**, and building more instrument than that is the classic way this +round turns into a project. Sample across artifact kinds rather than taking the first few hundred +rows of one file. + +Build a **definitions oracle the same way**, for 1.2(2) below: scan the raw bytes for `PROGRAM-ID`, +paragraph names and `SECTION` headers, hand-check a sample the same way, and use it to decide whether +a definition the stack reports is real. Nothing upstream of this prompt has built that oracle yet, so +it is the one instrument in Part 1 you are establishing from nothing rather than reproducing — state +your rule if you build it differently. -Say how many rows you checked and how you sampled them. A recall figure without its denominator and -its sampling rule is an impression. +Say how many rows you checked, for both oracles, and how you sampled them. A recall or +false-definition figure without its denominator and its sampling rule is an impression. ### 1.2 — the numbers @@ -143,6 +150,14 @@ Our gap list for COBOL is **guesswork about what real IBM i shops use**. It is: least one, per corpus. That reorders our roadmap directly, and no amount of reading the standard substitutes for it. +Count on the raw source text, not the parse tree: a case-insensitive literal match for each +construct's keyword or clause, skipping comment lines (column 7 holds `*` or `/` in fixed format) and +skipping matches inside string literals. Every match increments the occurrence count, including +repeats within one paragraph; a file counts once toward "files containing at least one" no matter how +many matches it holds. If your corpus needs a different rule — free format has no column 7 to skip, +for instance — state the rule you used instead; an unstated counting rule is what makes two runs of +this section incomparable. + Run this on the public corpus and expect the answer to be skewed: **public sample code under-represents CICS and embedded SQL badly**, because teaching material and demo applications are written to run without a transaction monitor or a database. Saying so with numbers is more useful @@ -230,8 +245,8 @@ retiring the fork sooner rather than later. One document, ordered: setup header, Part 1 numbers, Part 1b numbers if you ran it, the Part 2 answer. Then: -- **Every number carries its denominator and its instrument.** "98.7 / 100 on hand-checked rows - sampled across three corpora" is a result; "definitions are good" is not. +- **Every number carries its denominator and its instrument.** "96.7% of 716 files fully covered, + hand-checked against a lexical reference" is a result; "definitions are good" is not. - **A zero is a measurement; absent is not zero.** A count that cannot be a total is a floor and says so. Dynamic CALL is a floor. Bytes-not-parsed is a total. - **Report the losses in their own section**, not folded into an average. The corpus where the parse