Skip to content

ara-diac-small: full-set correction — 8.259 DER-CE; subset 3.658 unrepresentative - #54

Merged
ronaldtse merged 1 commit into
mainfrom
docs/ara-fullset
Aug 26, 2026
Merged

ara-diac-small: full-set correction — 8.259 DER-CE; subset 3.658 unrepresentative#54
ronaldtse merged 1 commit into
mainfrom
docs/ara-fullset

Conversation

@ronaldtse

Copy link
Copy Markdown
Contributor

Summary

  • Full-set SadeedDiac-25 measurement of ara-diac-small-1.0: 8.259 DER-CE on all 1,200 paragraphs vs 3.658 on the first-300 subset. Teacher reproduces 2.5815 against its documented 2.5793, confirming protocol consistency. The subset was not representative — it sits in the student's training-domain neighborhood; the remaining 900 paragraphs expose a domain-generalization gap
  • Catalog number is now the full-set figure (+5.68pp against the strict gate, disclosed); the leaderboard reading becomes "behind Sadeed-1.5B" (the earlier between-Gemini-Flash-and-GPT-4 placement was a subset artifact, withdrawn)
  • Standing rule recorded: student-tier numbers publish from full benchmark sets only; subset figures labeled at first publication
  • Paper integrity fix: frontier finding 3 now matches the results log — the Arabic 33M collapse ran on corrupted labels (retracted there), so the clean-label pretrained-or-collapse evidence is Thai; abstract/contribution/scope statements aligned
  • Metrics chain: one mapping entry with four table specs (subset + full-set) for fp32, plus a new entry covering the int8 variant's metadata; imf metrics passes for all 10 models

Test plan

  • imf cli metrics — all models trace to RESULTS.md (validated against the pushed branch ref; flip ref to main post-merge)
  • asciidoctor render clean, no stale 3.66/2.34 references outside the disclosed-context sentences
  • teacher reproduction within 0.003pp rules out a harness change

…58 was unrepresentative

Full 1,200-paragraph measurement (teacher reproduces 2.5815 vs
documented 2.5793, confirming protocol consistency): student scores
8.259 vs 3.658 on the first-300 subset. The subset sits in the
training-domain neighborhood; the remaining 900 paragraphs expose a
domain-generalization gap the subset hid. Catalog number is now the
full-set figure; the miss against the strict gate is +5.68pp,
disclosed. Standing rule: student-tier numbers publish from full sets
only, subset figures labeled at first publication.

Also corrects paper finding 3 to match the results log: the Arabic
33M collapse ran on corrupted labels (retracted there), so the
clean-label pretrained-or-collapse evidence is Thai.

Metrics chain: new ara-diac-small-1.0-fullset provenance entry
(der_teacher_fullset 2.5815 / der_student_fullset 8.259), metadata and
index updated for fp32 and int8 variants.
@ronaldtse
ronaldtse merged commit 0eeafd4 into main Aug 26, 2026
8 of 9 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant