feat(testcase): design validation and bug reports, plus a quick-start… - #22
Merged
Merged
Conversation
… table Two deliverables the skill was missing, added as references so SKILL.md stays small and neither loads unless the task needs it. design-validation.md: a Figma link or mockup is a requirement like any other. Discrepancies become D<n> rows in the same table, so step 6 lints design coverage the way it lints everything else. Compares tokens rather than rendered pixels — a hard-coded hex identical on screen to the theme variable is a finding a screenshot diff cannot see — and enumerates the states a frame hides: empty, loading, error, overflow, focus, dark mode, narrowest width. Draws the line where design authority stops, so placeholder copy and the gap between two provided breakpoints are not filed as bugs. bug-report.md: a report is a reproduction, not a notification. Exact values per step, expectation carrying its source, actual quoted, frequency as a ratio. Isolation before filing, and severity from user and data impact — with silent wrong data outranking a loud crash, and frequency kept out of severity. Quick start table routes the five entry points, and the workflow points at each reference at the step that needs it.
The reference had the content right and the operational shape wrong: a report that varies in layout cannot be imported into a tracker, and prose guidance on which environment details matter gets under-filled at the end of a shift. Fixes a hole the severity rules opened. Frequency was pushed out of severity with priority left to the team — and then nothing was handed to the team to set it with. User impact is now a required field, as is Type, which is what routes a security report to a different queue, and Workaround, which is what support needs on day one, before any fix. `None found` answers them; blank reads as not checked and the report comes back. Severity levels gain an example each — two readers agreeing on a criterion still disagree without one. Evidence gains screen recording, the only usable proof for a bug whose evidence is when things happened.
Self-review found the previous commit contradicting itself. Severity was listed above Workaround although S2 and S3 are told apart by it, so the field could not be answered in the order given. The Evidence row was left blank in the same table whose caption calls a blank field "not checked". `None found` was offered for Isolation, which is four checks with answers rather than something one finds. The commit message claimed Environment had been turned into a checklist; it had not. The table now carries each field's rule, so Reproduction, Evidence and Isolation stop being specified twice and drifting apart — only what a table cannot hold stays as prose. Type loses the `Data` value, undefined and overlapping Functional. The S4 example loses the clipped tooltip: unreadable content on a narrow screen is not cosmetic, and a miscalibrating example is worse than none. States plainly at the top that nothing lints any of this, unlike the case table. 107 lines to 89.
…at to the tracker Review raised two things the files asserted without earning. A design was treated as a requirement unconditionally. A Figma file also holds exploration, abandoned variants, stale components, placeholder copy, notes between designers and states nobody approved — and validating against the wrong frame does not yield a missed finding, it yields confident noise: a lint-clean, traceable D-row for a decision nobody made. A gate now runs first: exact node, version against the build under test, approval, mapping to a requirement ID, page conventions, and access to the values. Anything it blocks leaves as TBD rather than a finding. Values are read, never estimated. A hex sampled from a screenshot or a gap measured in pixels is a guess; with no access the pass reports structural findings only and says values were unverifiable. Accessibility checks each name the only source they may come from, and report as not verified when no tool or DOM is reachable — an invented contrast finding costs a developer a day and the next report its credibility. A section names the tools that measure better than this pass does, leaving it the judgement they cannot make. The bug report loses its field table. GitHub issue forms, Jira and Linear enforce required fields with a schema; a markdown template enforces nothing and drifts from what the team uses. What remains is what no form checks, and the severity model now says plainly that it is this skill's convention, that impact-times-likelihood is the common alternative, and that a project matrix wins.
…se for the lint Running summarize.py against the shape design-validation.md prescribed failed it: seven lint errors, missing ID, Category and Priority. The five-column design table was invented and does not parse, so the claim that step 6 checks design coverage like everything else was false at the parse step. Findings now use the step-3 table unchanged, with Req = D<n>. The same run showed every design row tripping success-path-only, because UI is not risk coverage and a visual finding attracts that category by reflex. Cases are now categorised by what they exercise: an empty state is a Boundary, a failed fetch an Error. State and UI are named as not counting, since that is what the script does. The gate splits into blocks, downgrades and limits, so a missing approval no longer stops a pass that a missing node genuinely must stop. Adds what the files did not answer: a design contradicting the written spec yields one TBD naming both sources, never a finding, with the spec winning behaviour and the design winning presentation; and native platforms read the accessibility tree or the framework's own props rather than reporting every check unverified.
Ran step 4 for real against a written spec. Both reviewers independently caught two cases whose expected deduction assumed weekends are not deductible, a rule the spec never states — on a table summarize.py had just passed as clean, which is the line between the two mechanisms: the script checks structure and coverage, only the second pass checks whether the number is right. They then disagreed on what the number should be, one recomputing four deductible days and the other writing three while listing four. Step 5 said to merge and drop overlap, and said nothing about a contradiction, so the obvious move is to take the more convincing reviewer — which is how a wrong value gets adopted with an independent pass cited behind it. The disagreement is now named as the signal that nobody derived the value, and the merge recomputes it from the requirement rather than choosing a side.
Two rounds of real second-pass review split on the same cell. One reviewer held that a spec excluding public holidays and saying nothing about weekends is not ambiguous, so TBD manufactures a question the requirement does not contain and ships a row that can never run. The other held that the spec never defines whether a day is a calendar day or a working day, so neither value is derivable. Both are right about a different failure: guessing invents a rule, and a bare TBD invents a doubt. Rule 5 now has three outcomes. Where the requirement is silent but one reading follows from the rules it does state, the case carries that value with the assumption named in the cell and the question raised in step 8 — runnable, falsifiable, and honest about what it assumed. TBD is reserved for a requirement that genuinely supports two readings giving different values.
…p the loop re-litigating one Four rounds of real review against one written spec. The same lens reversed itself three times on one cell: the literal value, then TBD, then the literal value again, each time argued from the same requirement. That is not a table defect the loop can grind out — it is the requirement owner's decision, and a fifth round buys a third answer rather than agreement. Such a cell is now frozen and escalated with both readings and the reversals recorded, and it no longer counts as blocking convergence. Rule 5's third branch needed its boundary drawn, because the reversals were about where it sits. Two readings qualify only when both have text behind them: 'the next anniversary' points grammatically at either anniversary, while a spec that excludes public holidays and never mentions weekends is silent, not ambiguous, and the second reading imports a working-day concept from outside. Over-hedging is named as its own failure, since a table half full of TBD cannot run and buries the questions it raises. Adds the lesson from a case written to be independent of an open question: it reached past two candidate expiry dates through a window wide enough for the next annual grant to land inside it, which a reviewer caught and the case's own note denied. A span chosen to outrun one rule must state what holds the others still.
…wn trigger Six rounds of independent review passed over two boundary cases built the usual way, by moving the join date back a day to sit just under a service tier. The seventh reviewer pointed out that one day before three years of service is not an anniversary, and the grant fires on anniversaries: the scenario cannot occur through the trigger at all, and the cases also contradicted the no-op-off-anniversary model another case in the same table establishes. Nudging an input by one unit is how a boundary usually gets built, so this is worth stating rather than leaving to judgement. The same off-by-one is caught by putting the boundary on a state the system does reach, the last anniversary below the threshold, and the case then describes a situation a user can actually be in. Preconditions are checked against the rules the other cases establish, since one that contradicts them is describing a state under test that cannot exist.
Three lessons with evidence from the trial run behind this branch. A precondition's numbers are derived, not free variables. The reachability rule was written about dates, and the next defect was a quantity: a fixture granting four days at an anniversary where the tier table mandates sixteen, which no correct implementation can be fixtured into. Both directions are now named, with the instruction to re-derive each quantity rather than read past it. Where an expected value depends on an unresolved question, an invariant true under every answer is usually available and is better than TBD: a balance that is zero under one reading and sixteen under the other is still never twenty, and twenty is what the bug produces. The case runs today and commits to nothing. Rule 5's third branch is for a value with no such invariant, not for every value an open question touches. A row an earlier round flagged is re-derived whole rather than at the patch. Reviewers stop at the first cause; the fix for it introduces the second in the same cell. Three consecutive rounds here found the new defect inside the line the previous fix had touched.
…case touches The invariant rule shipped with an example that does not satisfy it. A balance asserted as never twenty is invariant across the two readings of the expiry question, and not across the grant question sitting beside it: where a grant replaces rather than adds, the later grant overwrites the total to sixteen whether or not expiry ran, sixteen lies inside the accepted range, and the case stops telling a broken implementation from a correct one. The failure is quiet, which is what makes it worth a rule rather than a note. An inequality looks like it has already done the hedging, so the second open question inside it goes unlooked-for. Cases now list the questions they depend on before their inequality is trusted.
A question recorded as three live readings drew one case, separating the reading that felt most wrong from the other two. The remaining pair then passed every case in the table identically, so an implementation drifting between them was invisible — here, whether a request may be cancelled by its employee alone or also by their direct manager, with the direct manager never once placed in the cancelling position. The readings get counted, and each pair has to be told apart by some case.
…te a TBD Two ways a case can look sound and not be. A case that cannot tell the bug from correct behaviour invites a fix that invents the observable it needs. Here a re-approval guard moved no balance either way, so the assertion reached for an audit row and a timestamp the requirement never mentions — and a compliant implementation tracking only current status would fail it for a reason unrelated to the rule. Nothing catches this: the row traces to a requirement, names a wrong implementation and lints clean. Cases assert what the requirement defines, and where that cannot separate the behaviours they ask what the system records instead of deciding it. A TBD expected result is also never automatable. There is nothing to assert, and writing the test anyway hard-codes one reading of an open question into the suite — the guess rule 5 forbids, arriving through the metadata column rather than the cell.
…mitment A concurrency case declared itself neutral on an open question by citing the case beside it as already neutral. That sibling was not: the failure it was built to catch only occurs if the deduction happens at approval, one of the two answers the question is still open between. A third case in the same table assumed the other answer, so three rows encoded two mutually exclusive readings while the question's own ledger entry listed none of them. Neutrality is derived from the requirement rather than borrowed, and every case an open question touches is listed under that question — which is what makes two cases answering it in opposite directions visible instead of something a reviewer has to notice.
Every case for a rule that excludes something put the excluded item inside the input and checked it was skipped. None put one just outside and checked nothing changed, so an implementation scanning a window wider than the request would pass all of them while quietly discounting days the employee never asked for. The asymmetry is easy to reproduce: the inside case is what the requirement describes, and the outside case is what it implies. Rules that remove, filter or exempt get both.
A boundary case made reachable by shifting it onto a state the system actually produces can stop discriminating, while its claim about what it catches stays behind. Two tier-boundary cases moved off the threshold onto the previous anniversary still catch a service-year miscount, but at a value a full year below the tier both the greater-than and greater-or-equal implementations return the same number — so the operator bug they named, and still named afterwards, is caught only by the cases sitting exactly on the threshold. Distinguishes from is a claim about the input. When the input moves, the claim is re-derived and restated as what the case catches now.
…t names Distinguishes from is a claim, and it is the one part of a case nothing else checks: the lint sees a filled cell and a reviewer reading for coverage sees a plausible sentence. Two of them were false for ten rounds. A tier case sat a full year below its threshold, where the greater-than and greater-or-equal implementations return the same number. A balance case rejected its request under the correct arithmetic, under the double-subtracted holiday and under the gross range length alike, so all three were indistinguishable by its own assertion. The check is mechanical: construct the implementation the column names, run the case's stated input through it, and confirm the result differs. Where no input makes them diverge, the case is decoration.
Answers the review asking for evidence that these instructions produce better output, with the artefacts rather than a claim: the requirement, the table the skill produced against it, and the questions it could not answer without the requirement owner. The run is the argument. A twelve-case table passed the lint while two of its expected values assumed a rule the requirement never states, and both independent reviewers caught it. Coverage converged after four rounds; the fourteen that followed were spent on undisclosed assumptions, most of them introduced by the previous round's own fix. The loop stopped where the skill says to stop, with only labels left. The example also records what the loop cannot reach: twenty-six open questions remain, two of which dissolve depending on a third, and no further round closes any of them.
Late rounds of the trial produced rules at a cost the trial itself showed was not worth paying. Three go. The case built to outrun an open question, the hedge copied from a sibling row and the input moved to fix one property while breaking another were three descriptions of the same failure: a fix applied to one part of a row while the rest of it went unre-derived. Re-deriving a flagged row whole already says that, and says it once. The qualifier that an invariant must hold across every open question a case touches folds into the invariant rule as a clause, since it is a condition on that rule rather than a separate instruction. What stays is what the run showed earns its tokens: recomputing a numeric disagreement instead of picking a reviewer, freezing a cell reviewers reverse on, the line between silence and ambiguity, reachability of a boundary, invariants over TBD, not inventing a field to discriminate, and verifying a Distinguishes from claim by running the case's own input through it.
Measured both versions of this skill on one fresh logic-only spec: two agents wrote a first-pass table, each following one version and nothing else, and two blind judges audited both with the same standard. The result does not support keeping the added rules where they were. Defect counts came out level, and the table written under the longer version was the narrower one: it missed the requirement's largest ambiguity entirely without even logging it as a question, applied a rounding convention it never disclosed, and wrote a permission case that could not separate a UI-only check from a server-side one. Those are the three things its extra rules were about. Both tables also broke a rule the original already had, one requirement per Req cell, which is what 109 extra lines of reading tends to do. The rules came from review findings and read as questions asked of a finished table, so they go to the reviewer at step 4 where that is what they are. The authoring path returns to roughly its previous length.
…e the problem The A/B showed the longer version producing the narrower table, and the user's reading of why is the one the evidence supports: prescribing more turns the work into checklist-following, and a checklist only covers what its author already thought of. The skill's own opening asks for understanding, risk and systematic exploration, which is the part that generalises. So the nine rules do not move to the reviewer, they go. What survives is what changes behaviour in a direction the measurement did not contradict: rule 5, cut to one paragraph, which pushes back on guessing and on over-hedging alike; rule 8, cut to four lines, because an unreachable boundary tests a path no trigger reaches; and one lens row, Discrimination, since running a case's own input through the implementation its column names is the only check that column ever gets. SKILL.md ends at 228 lines against main's 205, and the 23 are mostly the two new features' wiring rather than instructions on how to think.
… to walk past a crash Measured this reference against its absence on a planted design review: an export whose only signed, pre-build frame was not the newest one, values the export itself hedged as estimates, a designer's note to a colleague, and placeholder copy. It earns its place. The arm carrying it picked the right baseline and said why, declined to file anything resting on an estimated colour or a frame edited after the build shipped, and demoted the pinned note to a question. The arm without it read the newest frame as canonical, quoted two hedged values as evidence, and turned the note into a medium-severity defect — two findings that cost a developer a day and fix nothing. Two things it got wrong go back in. Its state list read as items to file rather than questions to ask, so a screen no requirement asks to theme drew a missing-dark-mode finding; the list now says plainly that a state the product does not have is not a gap. And keeping to design scope let a call to an undefined function pass unremarked in the very file being read for tokens — scope stops you inventing findings, not reporting a crash.
Four lines where the rules around it run one or two, for a point that fits in three: a boundary nudged by one unit invents states the system never produces, and preconditions are derived rather than chosen.
The example was four files of a run whose own conclusion was that most of its eighteen rounds were spent repairing the previous round's patches. Three measured comparisons since then say more in a fraction of the space, including the one that removed nine of this branch's own rules. Trims design-validation's two longest sections: the write-up section restated the step-3 schema it had just cited, and the accessibility table listed each check on its own row where four groupings carry the same instruction.
Three arms, three reps each, on a spec line that supports two readings and code satisfying one of them. No guidance filed a confirmed bug every time, each rep inventing a different reading. A two-line compression of the rule halved that; the rule as written halved it again. Worth recording because the obvious edit was to shorten it, and the measurement said the opposite: a rule that picks a branch can be one line, a rule that states what an output must contain cannot.
…thing Three reps with the reference and three without, on one session's findings: a loud deterministic crash and a silently wrong total that was intermittent and took an afternoon to catch. Without the reference, all three runs rated the silent wrong total above the crash, one of them reasoning that failing loudly is what keeps the crash a notch below. With the reference and its four-level table, one run rated them equal and one inverted them — that run quoting the matrix as saying wrong money is S1 and then downgrading it anyway. A table is something to argue with; the judgement was already right without one. What the reference did earn, on the same runs: exact values in the steps 1 of 3 to 3 of 3, an expectation carrying its source 2 to 3, frequency as a ratio 2 to 3, and isolation — regression baseline, clean session, frontend versus backend — 0 of 3 to 3 of 3. Those are procedure nobody performs by default. Severity is judgement, and it needed one sentence, not a matrix.
…eport test design-validation pointed at a four-level scale in bug-report that no longer exists. The one thing it added to that scale is kept: an accessibility failure is never cosmetic. measured.md gains the fifth comparison, the one that removed the severity table and kept the isolation checklist.
…not a bug Five specs run end to end under both versions, judged blind. The newer one found 27 real defects to 11 and missed 1 to 17, for 15% more tokens. It also filed six false positives to one, and all six came from two lines this branch had added. One asked for any real defect in the code to be reported even outside design scope, and was read as licence to file robustness gaps no requirement mentions — a null user, a negative rate. The other called a hard-coded value that renders identically to its token a finding, which judges on two specs rejected because neither the requirement nor the frame asks for token binding. On the third spec the token finding stood, because there the value differed from the approved frame. Now: code contradicting a requirement is reported, wrong result or crash; a gap no requirement covers goes to questions; a value that differs from the frame is a finding and one that matches is a note. Re-measured with three independent agents each way, false positives fell from 2 of 3 to 0 of 3 with real defects 5.0 to 5.3 of 6. A first narrowing said crash instead of contradiction and looked like it had dropped a logic defect. Independent agents showed that loss, and two others a single agent reported, were noise; the within-agent repetitions used earlier correlate too strongly to count separately, which measured.md now says.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
… table
Two deliverables the skill was missing, added as references so SKILL.md stays small and neither loads unless the task needs it.
design-validation.md: a Figma link or mockup is a requirement like any other. Discrepancies become D rows in the same table, so step 6 lints design coverage the way it lints everything else. Compares tokens rather than rendered pixels — a hard-coded hex identical on screen to the theme variable is a finding a screenshot diff cannot see — and enumerates the states a frame hides: empty, loading, error, overflow, focus, dark mode, narrowest width. Draws the line where design authority stops, so placeholder copy and the gap between two provided breakpoints are not filed as bugs.
bug-report.md: a report is a reproduction, not a notification. Exact values per step, expectation carrying its source, actual quoted, frequency as a ratio. Isolation before filing, and severity from user and data impact — with silent wrong data outranking a loud crash, and frequency kept out of severity.
Quick start table routes the five entry points, and the workflow points at each reference at the step that needs it.