Skip to content

feat(testcase): design validation and bug reports, plus a quick-start… - #22

Merged
xtieume merged 28 commits into
mainfrom
skills/testcase
Sep 23, 2026
Merged

xtieume merged 28 commits into
mainfrom
skills/testcase

Conversation

@xtieume

@xtieume xtieume commented Sep 22, 2026

Copy link
Copy Markdown
Owner

… table

Two deliverables the skill was missing, added as references so SKILL.md stays small and neither loads unless the task needs it.

design-validation.md: a Figma link or mockup is a requirement like any other. Discrepancies become D rows in the same table, so step 6 lints design coverage the way it lints everything else. Compares tokens rather than rendered pixels — a hard-coded hex identical on screen to the theme variable is a finding a screenshot diff cannot see — and enumerates the states a frame hides: empty, loading, error, overflow, focus, dark mode, narrowest width. Draws the line where design authority stops, so placeholder copy and the gap between two provided breakpoints are not filed as bugs.

bug-report.md: a report is a reproduction, not a notification. Exact values per step, expectation carrying its source, actual quoted, frequency as a ratio. Isolation before filing, and severity from user and data impact — with silent wrong data outranking a loud crash, and frequency kept out of severity.

Quick start table routes the five entry points, and the workflow points at each reference at the step that needs it.

… table

Two deliverables the skill was missing, added as references so SKILL.md stays
small and neither loads unless the task needs it.

design-validation.md: a Figma link or mockup is a requirement like any other.
Discrepancies become D<n> rows in the same table, so step 6 lints design
coverage the way it lints everything else. Compares tokens rather than
rendered pixels — a hard-coded hex identical on screen to the theme variable
is a finding a screenshot diff cannot see — and enumerates the states a frame
hides: empty, loading, error, overflow, focus, dark mode, narrowest width.
Draws the line where design authority stops, so placeholder copy and the gap
between two provided breakpoints are not filed as bugs.

bug-report.md: a report is a reproduction, not a notification. Exact values
per step, expectation carrying its source, actual quoted, frequency as a
ratio. Isolation before filing, and severity from user and data impact — with
silent wrong data outranking a loud crash, and frequency kept out of severity.

Quick start table routes the five entry points, and the workflow points at
each reference at the step that needs it.
The reference had the content right and the operational shape wrong: a report
that varies in layout cannot be imported into a tracker, and prose guidance
on which environment details matter gets under-filled at the end of a shift.

Fixes a hole the severity rules opened. Frequency was pushed out of severity
with priority left to the team — and then nothing was handed to the team to
set it with. User impact is now a required field, as is Type, which is what
routes a security report to a different queue, and Workaround, which is what
support needs on day one, before any fix. `None found` answers them; blank
reads as not checked and the report comes back.

Severity levels gain an example each — two readers agreeing on a criterion
still disagree without one. Evidence gains screen recording, the only usable
proof for a bug whose evidence is when things happened.
Self-review found the previous commit contradicting itself. Severity was
listed above Workaround although S2 and S3 are told apart by it, so the field
could not be answered in the order given. The Evidence row was left blank in
the same table whose caption calls a blank field "not checked". `None found`
was offered for Isolation, which is four checks with answers rather than
something one finds. The commit message claimed Environment had been turned
into a checklist; it had not.

The table now carries each field's rule, so Reproduction, Evidence and
Isolation stop being specified twice and drifting apart — only what a table
cannot hold stays as prose. Type loses the `Data` value, undefined and
overlapping Functional. The S4 example loses the clipped tooltip: unreadable
content on a narrow screen is not cosmetic, and a miscalibrating example is
worse than none.

States plainly at the top that nothing lints any of this, unlike the case
table. 107 lines to 89.
…at to the tracker

Review raised two things the files asserted without earning.

A design was treated as a requirement unconditionally. A Figma file also holds
exploration, abandoned variants, stale components, placeholder copy, notes
between designers and states nobody approved — and validating against the
wrong frame does not yield a missed finding, it yields confident noise: a
lint-clean, traceable D-row for a decision nobody made. A gate now runs
first: exact node, version against the build under test, approval, mapping to
a requirement ID, page conventions, and access to the values. Anything it
blocks leaves as TBD rather than a finding.

Values are read, never estimated. A hex sampled from a screenshot or a gap
measured in pixels is a guess; with no access the pass reports structural
findings only and says values were unverifiable. Accessibility checks each
name the only source they may come from, and report as not verified when no
tool or DOM is reachable — an invented contrast finding costs a developer a
day and the next report its credibility. A section names the tools that
measure better than this pass does, leaving it the judgement they cannot make.

The bug report loses its field table. GitHub issue forms, Jira and Linear
enforce required fields with a schema; a markdown template enforces nothing
and drifts from what the team uses. What remains is what no form checks, and
the severity model now says plainly that it is this skill's convention, that
impact-times-likelihood is the common alternative, and that a project matrix
wins.
…se for the lint

Running summarize.py against the shape design-validation.md prescribed failed
it: seven lint errors, missing ID, Category and Priority. The five-column
design table was invented and does not parse, so the claim that step 6 checks
design coverage like everything else was false at the parse step. Findings now
use the step-3 table unchanged, with Req = D<n>.

The same run showed every design row tripping success-path-only, because UI is
not risk coverage and a visual finding attracts that category by reflex. Cases
are now categorised by what they exercise: an empty state is a Boundary, a
failed fetch an Error. State and UI are named as not counting, since that is
what the script does.

The gate splits into blocks, downgrades and limits, so a missing approval no
longer stops a pass that a missing node genuinely must stop. Adds what the
files did not answer: a design contradicting the written spec yields one TBD
naming both sources, never a finding, with the spec winning behaviour and the
design winning presentation; and native platforms read the accessibility tree
or the framework's own props rather than reporting every check unverified.
Ran step 4 for real against a written spec. Both reviewers independently
caught two cases whose expected deduction assumed weekends are not deductible,
a rule the spec never states — on a table summarize.py had just passed as
clean, which is the line between the two mechanisms: the script checks
structure and coverage, only the second pass checks whether the number is
right.

They then disagreed on what the number should be, one recomputing four
deductible days and the other writing three while listing four. Step 5 said to
merge and drop overlap, and said nothing about a contradiction, so the obvious
move is to take the more convincing reviewer — which is how a wrong value gets
adopted with an independent pass cited behind it. The disagreement is now
named as the signal that nobody derived the value, and the merge recomputes it
from the requirement rather than choosing a side.
Two rounds of real second-pass review split on the same cell. One reviewer
held that a spec excluding public holidays and saying nothing about weekends
is not ambiguous, so TBD manufactures a question the requirement does not
contain and ships a row that can never run. The other held that the spec never
defines whether a day is a calendar day or a working day, so neither value is
derivable. Both are right about a different failure: guessing invents a rule,
and a bare TBD invents a doubt.

Rule 5 now has three outcomes. Where the requirement is silent but one reading
follows from the rules it does state, the case carries that value with the
assumption named in the cell and the question raised in step 8 — runnable,
falsifiable, and honest about what it assumed. TBD is reserved for a
requirement that genuinely supports two readings giving different values.
…p the loop re-litigating one

Four rounds of real review against one written spec. The same lens reversed
itself three times on one cell: the literal value, then TBD, then the literal
value again, each time argued from the same requirement. That is not a table
defect the loop can grind out — it is the requirement owner's decision, and a
fifth round buys a third answer rather than agreement. Such a cell is now
frozen and escalated with both readings and the reversals recorded, and it no
longer counts as blocking convergence.

Rule 5's third branch needed its boundary drawn, because the reversals were
about where it sits. Two readings qualify only when both have text behind
them: 'the next anniversary' points grammatically at either anniversary, while
a spec that excludes public holidays and never mentions weekends is silent,
not ambiguous, and the second reading imports a working-day concept from
outside. Over-hedging is named as its own failure, since a table half full of
TBD cannot run and buries the questions it raises.

Adds the lesson from a case written to be independent of an open question: it
reached past two candidate expiry dates through a window wide enough for the
next annual grant to land inside it, which a reviewer caught and the case's own
note denied. A span chosen to outrun one rule must state what holds the others
still.
…wn trigger

Six rounds of independent review passed over two boundary cases built the
usual way, by moving the join date back a day to sit just under a service
tier. The seventh reviewer pointed out that one day before three years of
service is not an anniversary, and the grant fires on anniversaries: the
scenario cannot occur through the trigger at all, and the cases also
contradicted the no-op-off-anniversary model another case in the same table
establishes.

Nudging an input by one unit is how a boundary usually gets built, so this is
worth stating rather than leaving to judgement. The same off-by-one is caught
by putting the boundary on a state the system does reach, the last anniversary
below the threshold, and the case then describes a situation a user can
actually be in. Preconditions are checked against the rules the other cases
establish, since one that contradicts them is describing a state under test
that cannot exist.
Three lessons with evidence from the trial run behind this branch.

A precondition's numbers are derived, not free variables. The reachability
rule was written about dates, and the next defect was a quantity: a fixture
granting four days at an anniversary where the tier table mandates sixteen,
which no correct implementation can be fixtured into. Both directions are now
named, with the instruction to re-derive each quantity rather than read past
it.

Where an expected value depends on an unresolved question, an invariant true
under every answer is usually available and is better than TBD: a balance that
is zero under one reading and sixteen under the other is still never twenty,
and twenty is what the bug produces. The case runs today and commits to
nothing. Rule 5's third branch is for a value with no such invariant, not for
every value an open question touches.

A row an earlier round flagged is re-derived whole rather than at the patch.
Reviewers stop at the first cause; the fix for it introduces the second in the
same cell. Three consecutive rounds here found the new defect inside the line
the previous fix had touched.
…case touches

The invariant rule shipped with an example that does not satisfy it. A balance
asserted as never twenty is invariant across the two readings of the expiry
question, and not across the grant question sitting beside it: where a grant
replaces rather than adds, the later grant overwrites the total to sixteen
whether or not expiry ran, sixteen lies inside the accepted range, and the case
stops telling a broken implementation from a correct one.

The failure is quiet, which is what makes it worth a rule rather than a note.
An inequality looks like it has already done the hedging, so the second open
question inside it goes unlooked-for. Cases now list the questions they depend
on before their inequality is trusted.
A question recorded as three live readings drew one case, separating the
reading that felt most wrong from the other two. The remaining pair then
passed every case in the table identically, so an implementation drifting
between them was invisible — here, whether a request may be cancelled by its
employee alone or also by their direct manager, with the direct manager never
once placed in the cancelling position.

The readings get counted, and each pair has to be told apart by some case.
…te a TBD

Two ways a case can look sound and not be.

A case that cannot tell the bug from correct behaviour invites a fix that
invents the observable it needs. Here a re-approval guard moved no balance
either way, so the assertion reached for an audit row and a timestamp the
requirement never mentions — and a compliant implementation tracking only
current status would fail it for a reason unrelated to the rule. Nothing
catches this: the row traces to a requirement, names a wrong implementation
and lints clean. Cases assert what the requirement defines, and where that
cannot separate the behaviours they ask what the system records instead of
deciding it.

A TBD expected result is also never automatable. There is nothing to assert,
and writing the test anyway hard-codes one reading of an open question into
the suite — the guess rule 5 forbids, arriving through the metadata column
rather than the cell.
…mitment

A concurrency case declared itself neutral on an open question by citing the
case beside it as already neutral. That sibling was not: the failure it was
built to catch only occurs if the deduction happens at approval, one of the
two answers the question is still open between. A third case in the same table
assumed the other answer, so three rows encoded two mutually exclusive
readings while the question's own ledger entry listed none of them.

Neutrality is derived from the requirement rather than borrowed, and every
case an open question touches is listed under that question — which is what
makes two cases answering it in opposite directions visible instead of
something a reviewer has to notice.
Every case for a rule that excludes something put the excluded item inside the
input and checked it was skipped. None put one just outside and checked
nothing changed, so an implementation scanning a window wider than the request
would pass all of them while quietly discounting days the employee never asked
for.

The asymmetry is easy to reproduce: the inside case is what the requirement
describes, and the outside case is what it implies. Rules that remove, filter
or exempt get both.
A boundary case made reachable by shifting it onto a state the system actually
produces can stop discriminating, while its claim about what it catches stays
behind. Two tier-boundary cases moved off the threshold onto the previous
anniversary still catch a service-year miscount, but at a value a full year
below the tier both the greater-than and greater-or-equal implementations
return the same number — so the operator bug they named, and still named
afterwards, is caught only by the cases sitting exactly on the threshold.

Distinguishes from is a claim about the input. When the input moves, the claim
is re-derived and restated as what the case catches now.
…t names

Distinguishes from is a claim, and it is the one part of a case nothing else
checks: the lint sees a filled cell and a reviewer reading for coverage sees a
plausible sentence. Two of them were false for ten rounds. A tier case sat a
full year below its threshold, where the greater-than and greater-or-equal
implementations return the same number. A balance case rejected its request
under the correct arithmetic, under the double-subtracted holiday and under
the gross range length alike, so all three were indistinguishable by its own
assertion.

The check is mechanical: construct the implementation the column names, run
the case's stated input through it, and confirm the result differs. Where no
input makes them diverge, the case is decoration.
Answers the review asking for evidence that these instructions produce better
output, with the artefacts rather than a claim: the requirement, the table the
skill produced against it, and the questions it could not answer without the
requirement owner.

The run is the argument. A twelve-case table passed the lint while two of its
expected values assumed a rule the requirement never states, and both
independent reviewers caught it. Coverage converged after four rounds; the
fourteen that followed were spent on undisclosed assumptions, most of them
introduced by the previous round's own fix. The loop stopped where the skill
says to stop, with only labels left.

The example also records what the loop cannot reach: twenty-six open questions
remain, two of which dissolve depending on a third, and no further round
closes any of them.
Late rounds of the trial produced rules at a cost the trial itself showed was
not worth paying. Three go.

The case built to outrun an open question, the hedge copied from a sibling row
and the input moved to fix one property while breaking another were three
descriptions of the same failure: a fix applied to one part of a row while the
rest of it went unre-derived. Re-deriving a flagged row whole already says
that, and says it once.

The qualifier that an invariant must hold across every open question a case
touches folds into the invariant rule as a clause, since it is a condition on
that rule rather than a separate instruction.

What stays is what the run showed earns its tokens: recomputing a numeric
disagreement instead of picking a reviewer, freezing a cell reviewers reverse
on, the line between silence and ambiguity, reachability of a boundary,
invariants over TBD, not inventing a field to discriminate, and verifying a
Distinguishes from claim by running the case's own input through it.
Measured both versions of this skill on one fresh logic-only spec: two agents
wrote a first-pass table, each following one version and nothing else, and two
blind judges audited both with the same standard.

The result does not support keeping the added rules where they were. Defect
counts came out level, and the table written under the longer version was the
narrower one: it missed the requirement's largest ambiguity entirely without
even logging it as a question, applied a rounding convention it never
disclosed, and wrote a permission case that could not separate a UI-only check
from a server-side one. Those are the three things its extra rules were about.
Both tables also broke a rule the original already had, one requirement per
Req cell, which is what 109 extra lines of reading tends to do.

The rules came from review findings and read as questions asked of a finished
table, so they go to the reviewer at step 4 where that is what they are. The
authoring path returns to roughly its previous length.
…e the problem

The A/B showed the longer version producing the narrower table, and the user's
reading of why is the one the evidence supports: prescribing more turns the
work into checklist-following, and a checklist only covers what its author
already thought of. The skill's own opening asks for understanding, risk and
systematic exploration, which is the part that generalises.

So the nine rules do not move to the reviewer, they go. What survives is what
changes behaviour in a direction the measurement did not contradict: rule 5,
cut to one paragraph, which pushes back on guessing and on over-hedging alike;
rule 8, cut to four lines, because an unreachable boundary tests a path no
trigger reaches; and one lens row, Discrimination, since running a case's own
input through the implementation its column names is the only check that
column ever gets.

SKILL.md ends at 228 lines against main's 205, and the 23 are mostly the two
new features' wiring rather than instructions on how to think.
… to walk past a crash

Measured this reference against its absence on a planted design review: an
export whose only signed, pre-build frame was not the newest one, values the
export itself hedged as estimates, a designer's note to a colleague, and
placeholder copy.

It earns its place. The arm carrying it picked the right baseline and said
why, declined to file anything resting on an estimated colour or a frame
edited after the build shipped, and demoted the pinned note to a question.
The arm without it read the newest frame as canonical, quoted two hedged
values as evidence, and turned the note into a medium-severity defect — two
findings that cost a developer a day and fix nothing.

Two things it got wrong go back in. Its state list read as items to file
rather than questions to ask, so a screen no requirement asks to theme drew a
missing-dark-mode finding; the list now says plainly that a state the product
does not have is not a gap. And keeping to design scope let a call to an
undefined function pass unremarked in the very file being read for tokens —
scope stops you inventing findings, not reporting a crash.
Four lines where the rules around it run one or two, for a point that fits in
three: a boundary nudged by one unit invents states the system never produces,
and preconditions are derived rather than chosen.
The example was four files of a run whose own conclusion was that most of its
eighteen rounds were spent repairing the previous round's patches. Three
measured comparisons since then say more in a fraction of the space, including
the one that removed nine of this branch's own rules.

Trims design-validation's two longest sections: the write-up section restated
the step-3 schema it had just cited, and the accessibility table listed each
check on its own row where four groupings carry the same instruction.
Three arms, three reps each, on a spec line that supports two readings and
code satisfying one of them. No guidance filed a confirmed bug every time,
each rep inventing a different reading. A two-line compression of the rule
halved that; the rule as written halved it again.

Worth recording because the obvious edit was to shorten it, and the
measurement said the opposite: a rule that picks a branch can be one line, a
rule that states what an output must contain cannot.
…thing

Three reps with the reference and three without, on one session's findings: a
loud deterministic crash and a silently wrong total that was intermittent and
took an afternoon to catch.

Without the reference, all three runs rated the silent wrong total above the
crash, one of them reasoning that failing loudly is what keeps the crash a
notch below. With the reference and its four-level table, one run rated them
equal and one inverted them — that run quoting the matrix as saying wrong
money is S1 and then downgrading it anyway. A table is something to argue
with; the judgement was already right without one.

What the reference did earn, on the same runs: exact values in the steps 1 of
3 to 3 of 3, an expectation carrying its source 2 to 3, frequency as a ratio 2
to 3, and isolation — regression baseline, clean session, frontend versus
backend — 0 of 3 to 3 of 3. Those are procedure nobody performs by default.
Severity is judgement, and it needed one sentence, not a matrix.
…eport test

design-validation pointed at a four-level scale in bug-report that no longer
exists. The one thing it added to that scale is kept: an accessibility failure
is never cosmetic.

measured.md gains the fifth comparison, the one that removed the severity
table and kept the isolation checklist.
…not a bug

Five specs run end to end under both versions, judged blind. The newer one
found 27 real defects to 11 and missed 1 to 17, for 15% more tokens. It also
filed six false positives to one, and all six came from two lines this branch
had added.

One asked for any real defect in the code to be reported even outside design
scope, and was read as licence to file robustness gaps no requirement mentions
— a null user, a negative rate. The other called a hard-coded value that
renders identically to its token a finding, which judges on two specs rejected
because neither the requirement nor the frame asks for token binding. On the
third spec the token finding stood, because there the value differed from the
approved frame.

Now: code contradicting a requirement is reported, wrong result or crash; a
gap no requirement covers goes to questions; a value that differs from the
frame is a finding and one that matches is a note. Re-measured with three
independent agents each way, false positives fell from 2 of 3 to 0 of 3 with
real defects 5.0 to 5.3 of 6.

A first narrowing said crash instead of contradiction and looked like it had
dropped a logic defect. Independent agents showed that loss, and two others a
single agent reported, were noise; the within-agent repetitions used earlier
correlate too strongly to count separately, which measured.md now says.
@xtieume
xtieume merged commit ed0aea4 into main Sep 23, 2026
2 of 3 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant