From 1016af1e95c76f63e2e5dfdf6083f9a880a0619a Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Fri, 28 Aug 2026 12:52:55 +0800 Subject: [PATCH 1/6] =?UTF-8?q?docs:=20r7=20sweeps=20ID+OOD=20=E2=80=94=20?= =?UTF-8?q?new=20canonical=20Arabic=20teacher;=20paper=20leaderboard=20to?= =?UTF-8?q?=202.2864?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit WikiNews-2024 multi-ref: 17.3794/11.8273 vs r6's 19.8191/12.4613. In-domain 2.2864 vs 2.5793. The teacher lineage no longer trades ID for OOD. Future distillations take r7 as teacher. --- docs/PUBLICATION-NOTES.md | 18 ++++++++++-------- docs/paper.adoc | 7 ++++--- 2 files changed, 14 insertions(+), 11 deletions(-) diff --git a/docs/PUBLICATION-NOTES.md b/docs/PUBLICATION-NOTES.md index 3bdbf05..6f09bf8 100644 --- a/docs/PUBLICATION-NOTES.md +++ b/docs/PUBLICATION-NOTES.md @@ -83,14 +83,16 @@ the paper — first controlled Muon measurement on byte-level seq2seq distillation. Paper: frontier section. EXPERIMENTS.md E3, RESULTS.md run-004. Factorial cell 4 (vanilla+Muon, run-005) queued. -### 10. Arabic news-domain adaptation (r7) — ID LANDED 2026-08-28 -**2.2864 / 1.3343 full-set windowed zero-skip — new best dedicated -model, −0.29pp over r6** (the news mix improved in-domain, not just -OOD; behind only Claude-3.7-Sonnet's published 1.3941). OOD half -(WikiNews-2024 multi-ref, gate: beat r6's 19.82/12.46) running via the -auto-launched actor. → rababa RESULTS.md (recorded) → paper leaderboard -table + OOD note once OOD confirms; canonical-teacher promotion -pending that check. +### 10. Arabic news-domain adaptation (r7) — COMPLETE 2026-08-28, NEW CANONICAL TEACHER +**ID 2.2864/1.3343 (−0.29pp over r6) AND OOD WikiNews-2024 multi-ref +17.38/11.83 (vs r6's 19.82/12.46)** — r7 sweeps both surfaces; the +teacher-lineage domain trade-off is gone. Best dedicated model under +the protocol, behind only Claude-3.7-Sonnet's published 1.3941, now +well clear of GLM-5.2 (2.6911). r7 replaces r6 as the canonical +Arabic teacher for future distillations. Paper leaderboard table +updated; rababa RESULTS.md carries both tables. Follow-on (user +decision): ara-diac-2.0 teacher release; students re-distilled from +r7 with Muon (E3) are the natural ara-diac-small-2.0. ## Venue fit (working notes) diff --git a/docs/paper.adoc b/docs/paper.adoc index 836caeb..0eb1735 100644 --- a/docs/paper.adoc +++ b/docs/paper.adoc @@ -203,7 +203,8 @@ External claims are made only where a public benchmark exists and the full proto |System |Params |DER-CE |Claude 3.7 Sonnet (published) |— |1.3941 -|**our teacher (r6, ByT5-base)** |580M |**2.5793** +|**our teacher (r7, ByT5-base)** |580M |**2.2864** +|our previous teacher (r6) |580M |2.5793 |GLM-5.2 (reproduced, zero-skip) |— |2.6911 |Gemini Flash 2.0 (published) |— |3.1926 |GPT-4 (published) |— |3.8645 @@ -211,13 +212,13 @@ External claims are made only where a public benchmark exists and the full proto |our client student (ByT5-small) |300M |8.259 |=== -The teacher is the best dedicated (task-trained, runnable-locally) model measured under this protocol — second only to a frontier proprietary LLM's published number, ahead of a clean GLM-5.2 reproduction, Gemini Flash, and GPT-4 — and it does so at 580M parameters against Sadeed's 1.5B. +The teacher is the best dedicated (task-trained, runnable-locally) model measured under this protocol — second only to a frontier proprietary LLM's published number, ahead of a clean GLM-5.2 reproduction, Gemini Flash, and GPT-4 — and it does so at 580M parameters against Sadeed's 1.5B. The current rung (r7) adds a teacher-labeled news-domain mix to the morphological-auxiliary teacher (r6) and improves both surfaces at once: in-domain 2.5793 → 2.2864, and out-of-domain WikiNews-2024 multi-reference 19.82/12.46 → 17.38/11.83 (WER/DER) — the paragraph-specialization trade-off of the earlier lineage is gone. The r6 teacher's auxiliary-task design is validated by a controlled ablation: an identical run differing only in the auxiliary stream's output representation — broad-phonemic IPA of the same training units (deterministic converter) in place of morphological analysis — scores 2.6588 (vs 2.6775 for no auxiliary task and 2.5793 for the morphological one), while a held-out probe confirms the phonemic projection itself was learned (2.3% CER on the IPA stream). The improvement is therefore attributable to lexical-morphological knowledge injection, not to phonemic supervision per se: the strong form of the "phonemes help diacritization" hypothesis fails where its weak form (any structured auxiliary projection beats none) barely holds. The client student's row carries a measurement lesson we record rather than bury. Its first-published figure, 3.66, was measured on the first 300 paragraphs of the benchmark; the full 1,200-paragraph run scores 8.26 (teacher reproduces at 2.5815 against its documented 2.5793, confirming protocol consistency). The subset was not representative: its paragraphs sit closer to the student's training domain, and the remaining 900 expose a domain-generalization gap the subset hid. Two practices follow, now standing rules: student-tier numbers are published only from full benchmark sets, and a subset figure is labeled as such at first publication. The student ships with the full-set number disclosed — behind Sadeed-1.5B on the leaderboard — and closing its domain gap (more diverse label coverage, or on-policy distillation; <>) is open work. -Two protocol notes the benchmark's own authors make and we adopt: self-reported numbers from other protocols (Sadeed's repository reports 1.2% DER under its own split) are not comparable, and multi-reference evaluation (Mohamed & Mubarak 2025) is the honest treatment of valid-alternative diacritizations; on WikiNews-2024 our teacher family scores 19.82 WER against QCRI's in-domain-trained 2.70 — a domain gap (classical hadith vs news text), recorded rather than averaged away. +Two protocol notes the benchmark's own authors make and we adopt: self-reported numbers from other protocols (Sadeed's repository reports 1.2% DER under its own split) are not comparable, and multi-reference evaluation (Mohamed & Mubarak 2025) is the honest treatment of valid-alternative diacritizations; on WikiNews-2024 our teacher scores 17.38 WER against QCRI's in-domain-trained 2.70 — a domain gap (classical hadith vs news text), recorded rather than averaged away. *Hebrew.* No external leaderboard exists for the Nakdimon test domain (Biblical/Rabbinic). On that split our best teacher scores 16.58% DER — against DictaBERT-large, the reference production diacritizer, at 35.63% *on the same test, run by us*. Recent open systems (D-Nikud; visual-representation approaches, Elboher & Pinter 2025) report state of the art on *modern*-text benchmarks we do not evaluate. Hebrew remains the campaign's disclosed weak spot in absolute terms while being locally competitive; closing the domain gap is active work. From cc3bddfeed992953f749ad7701108767d8d3010d Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Fri, 28 Aug 2026 14:50:33 +0800 Subject: [PATCH 2/6] feat: ara-diac-2.0 release files + ara-diac-small-2.0 distill spec (E4) MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit ara-diac-2.0 = the r7 canonical teacher (ID 2.2864/1.3343, OOD 17.3794/11.8273 — full sweep over 1.0). Metadata + model card + export entry; provenance anchors in rababa RESULTS.md (r7 verdict tables). Export runs the new gates: int8 with fp32 head, CER parity + margin budget. ara-diac-small-2 = run-006-r7-muon: r7 teacher + Muon (the two measured wins compound). E4 pre-registered: gate ≤ 6.26 (≥2pp over the shipped 8.259), prediction 4.3-5.0. Orchestrator arm added; idempotent. --- docs/EXPERIMENTS.md | 16 +++++++++ models/ara-diac/ara-diac-2.0.README.md | 39 ++++++++++++++++++++++ models/ara-diac/ara-diac-2.0.metadata.yaml | 33 ++++++++++++++++++ src/gpu/modal_distill.py | 20 +++++++++++ src/gpu/modal_export.py | 9 +++++ 5 files changed, 117 insertions(+) create mode 100644 models/ara-diac/ara-diac-2.0.README.md create mode 100644 models/ara-diac/ara-diac-2.0.metadata.yaml diff --git a/docs/EXPERIMENTS.md b/docs/EXPERIMENTS.md index eae0c44..4c02642 100644 --- a/docs/EXPERIMENTS.md +++ b/docs/EXPERIMENTS.md @@ -143,6 +143,22 @@ All rows passed the CER parity gate at release. Readings: - **Report:** either direction goes to the paper's training-methods appendix. +## E4 — ara-diac-small-2.0 candidate (run-006-r7-muon) + +- **Status:** queued (registered 2026-08-28 before launch). +- **Hypothesis:** the two measured wins compound — the r7 canonical + teacher (better labels; ID 2.2864 vs 2.5793) plus the E3-adopted + Muon optimizer (−2.727pp on r6 labels) — moving the client rung far + below the shipped 8.259. +- **Design:** identical corpus/limits/schedule to run-002/003/004/005; + teacher = run-007-news/best (fresh greedy labels, teacher_labels_r7); + optimizer = Muon (embedding-like group on AdamW); vanilla ByT5-small + (the shipped architecture — the 2.0 ships what was measured). +- **Pre-agreed gate:** beats the shipped student by ≥2.0pp full-set + windowed DER (i.e. ≤ 6.26) to justify the 2.0 release; the strict + teacher+0.5pp budget remains the disclosed north star. +- **Prediction (registered):** 4.3–5.0, by E3's 4.829 on weaker labels. + ## Parked - **Speculative decoding** (LongCat converts sparsity→speed): revisit diff --git a/models/ara-diac/ara-diac-2.0.README.md b/models/ara-diac/ara-diac-2.0.README.md new file mode 100644 index 0000000..3b72cf1 --- /dev/null +++ b/models/ara-diac/ara-diac-2.0.README.md @@ -0,0 +1,39 @@ +# ara-diac-2.0 + +Arabic diacritization (haraqat restoration), ByT5-base, 580M parameters. + +## What changed from 1.0 + +ara-diac-1.0 (r6) was trained with a morphological auxiliary task +(qalsadi iʿrāb labels) on top of the paragraph-context r5 lineage. +ara-diac-2.0 (r7) adds a news-domain adaptation stage: 13,986 +teacher-labeled news units mixed at 0.85% with the r5-units anchor, +plus 400 gold WikiNews-2014 lines, initialized from r6. + +The result is a full sweep over 1.0 — both surfaces improved: + +| Surface | ara-diac-1.0 (r6) | ara-diac-2.0 (r7) | +|---|---|---| +| SadeedDiac-25, full 1,200, windowed zero-skip, greedy | DER 2.5793 / Morph 1.5317 | **DER 2.2864 / Morph 1.3343** | +| WikiNews-2024 multi-ref (QCRI protocol, full) | WER 19.8191 / DER 12.4613 | **WER 17.3794 / DER 11.8273** | + +The earlier lineage's in-domain/out-of-domain trade-off is gone: the +news mix improved in-domain substantially, not just OOD. + +## Positioning + +Best dedicated (task-trained, runnable-locally) model measured on +SadeedDiac-25 under the zero-skip Misraj protocol — behind only +Claude-3.7-Sonnet's published 1.3941, ahead of GLM-5.2 (2.6911 +zero-skip reproduction), Gemini-Flash-2.0 (3.1926), and GPT-4 (3.8645), +at 580M parameters against Sadeed-1.5B's 7.2915. + +## Contract + +IMF v1 zip: byte tokenizer (id = byte + 3, trailing EOS), KV-cache +decoder, opset 14. Greedy decode is the runtime protocol. int8 zips +keep the tied head in fp32 (the margin-gate fix); every precision is +gated by CER parity AND the argmax-flip margin budget. + +Provenance: rababa `train_arabic_r7.py`, run `rababa_arabic_byt5/run-007-news`; +verdict tables in `rababa/docs/RESULTS.md` (r7 sections). BSD-3-Clause. diff --git a/models/ara-diac/ara-diac-2.0.metadata.yaml b/models/ara-diac/ara-diac-2.0.metadata.yaml new file mode 100644 index 0000000..8a178b9 --- /dev/null +++ b/models/ara-diac/ara-diac-2.0.metadata.yaml @@ -0,0 +1,33 @@ +format: imf-v1 +id: ara-diac-2.0 +task: diacritization +source_script: Arab +target: Arab +tokenizer: bytes +opset: 14 +decoder: kv +precision: fp32 +license: BSD-3-Clause +trained_from: 'rababa train_arabic_r7.py run-007-news (news-domain adaptation of the r6 + morphological-aux teacher: r5-units anchor + 13,986 teacher-labeled news units at + 0.85% mix + 400 gold WikiNews-2014 lines, init from run-006-morph); checkpoint + rababa-checkpoints:/rababa_arabic_byt5/run-007-news/best' +metrics: +- name: der_total_greedy + value: 2.2864 + protocol: greedy decode (the v1 runtime path); SadeedDiac-25, windowed zero-skip + at 1400 bytes, full 1,200 paragraphs; ByT5-base r7 + source: rababa/docs/RESULTS.md#r7-verdict-table-sadeeddiac-25-2026-08-28 +- name: der_morph_greedy + value: 1.3343 + protocol: same harness; Morphological DER (word-final case endings only) + source: rababa/docs/RESULTS.md#r7-verdict-table-sadeeddiac-25-2026-08-28 +- name: wer_wikinews_multiref + value: 17.3794 + protocol: WikiNews-2024 multi-reference, QCRI EvalDiac protocol (full mode), 356 + texts / 10,616 words + source: rababa/docs/RESULTS.md#r7-ood-verdict-table-wikinews-2024-multiref-2026-08-28 +- name: der_wikinews_multiref + value: 11.8273 + protocol: same harness + source: rababa/docs/RESULTS.md#r7-ood-verdict-table-wikinews-2024-multiref-2026-08-28 diff --git a/src/gpu/modal_distill.py b/src/gpu/modal_distill.py index 547ffd4..dfccb71 100644 --- a/src/gpu/modal_distill.py +++ b/src/gpu/modal_distill.py @@ -205,6 +205,25 @@ def _ensure_src_path() -> None: "mode": "sequence", "note": "vanilla ByT5-small + Muon (factorial cell 4)", }, + "ara-diac-small-2": { + # ara-diac-small-2.0 candidate: the r7 canonical teacher + the + # E3-adopted Muon optimizer. Fresh labels from r7 (the v2 labels + # are r6-teacher). EXPERIMENTS.md E4. + "teacher": "rababa_arabic_byt5/run-007-news/best", + "teacher_volume": "rababa", + "student_init": "google/byt5-small", + "optimizer": "muon", + "muon_lr": "0.01", + "train": "r5-units/domain.txt", + "train_extra": ["r5-units/replay.txt"], + "unit_limits": [24000, 6000], + "max_len": 1450, + "label_beams": "1", + "out": "rababa_arabic_distill_small/run-006-r7-muon", + "labels_file": "teacher_labels_r7.jsonl", + "mode": "sequence", + "note": "r7 teacher + Muon; E4 gate: beat the shipped 8.259 by >= 2pp", + }, "ara-diac-tiny": { "teacher": "rababa_arabic_byt5/run-006-morph/best", "teacher_volume": "rababa", @@ -1672,6 +1691,7 @@ def qwen_next_chain() -> dict: ("ara-diac-small-pkm", "rababa_arabic_distill_small/run-003-pkm"), ("ara-diac-small-pkm-muon", "rababa_arabic_distill_small/run-004-pkm-muon"), ("ara-diac-small-muon", "rababa_arabic_distill_small/run-005-muon"), + ("ara-diac-small-2", "rababa_arabic_distill_small/run-006-r7-muon"), ] ROOT = Path("/checkpoints") diff --git a/src/gpu/modal_export.py b/src/gpu/modal_export.py index 75d29c8..cdefffb 100644 --- a/src/gpu/modal_export.py +++ b/src/gpu/modal_export.py @@ -87,6 +87,15 @@ "test_data": "nakdimon/test-imf.jsonl", "probe": "שלום", }, + "ara-diac2": { + "volume": "/volumes/rababa-checkpoints", + "checkpoint": "rababa_arabic_byt5/run-007-news/best", + "metadata": "models/ara-diac/ara-diac-2.0.metadata.yaml", + "readme": "models/ara-diac/ara-diac-2.0.README.md", + "test_volume": "/datasets/rababa", + "test_data": "arabic-sadeed-imf/test.jsonl", + "probe": "مكتبة", + }, "ara-diac": { "volume": "/volumes/rababa-checkpoints", "checkpoint": "rababa_arabic_byt5/run-006-morph/best", From 29007c10300b054a5d5aecdd8abdbefd682ca0ec Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Fri, 28 Aug 2026 21:57:52 +0800 Subject: [PATCH 3/6] =?UTF-8?q?docs:=20factorial=20cell=204=20=E2=80=94=20?= =?UTF-8?q?vanilla+Muon=205.2945;=20paper=20carries=20the=20full=202x2?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Optimizer alone −2.96pp, memory alone −0.70pp (−0.47 under Muon), combined −3.43pp — roughly additive. Optimization is the dominant recoverable term of the distillation gap; capacity second; ~2.2pp residual is domain coverage. --- docs/EXPERIMENTS.md | 19 +++++++++++++------ docs/paper.adoc | 14 ++++++++------ 2 files changed, 21 insertions(+), 12 deletions(-) diff --git a/docs/EXPERIMENTS.md b/docs/EXPERIMENTS.md index 4c02642..0eca8e0 100644 --- a/docs/EXPERIMENTS.md +++ b/docs/EXPERIMENTS.md @@ -123,12 +123,19 @@ All rows passed the CER parity gate at release. Readings: equal steps; step time ~1.2s vs ~3.4s (the <15% overhead gate met with margin — Newton–Schulz is cheap next to 1450-byte windows); no stability events. -- **Combined with E2, the gap decomposition at the ByT5-small rung:** - capacity ~0.70pp (PKM) + optimization ~2.73pp (Muon, on the PKM - architecture) + residual ~2.25pp (domain coverage — the subset - lesson). Caveat carried: the A/B isolates the optimizer ON the - memory architecture; the vanilla+Muon cell (run-005-muon) is queued - to complete the 2x2 factorial. +- **Factorial cell 4 (run-005-muon, vanilla + Muon, 2026-08-28): + 5.2945.** The 2x2 closes cleanly: + + | Full-set DER | AdamW | Muon | + |---|---|---| + | vanilla ByT5-small | 8.259 | 5.295 | + | + PKM memory | 7.555 | **4.829** | + + Decomposition (teacher ~2.6): optimizer alone −2.96pp; memory alone + −0.70pp (−0.47 under Muon); combined −3.43pp — roughly additive, + slightly sub-additive on the memory term. Optimization is the + dominant recoverable term at this rung; capacity is second; residual + ~2.2pp is domain coverage (the subset lesson). - **Hypothesis:** orthogonalized-momentum updates (Newton–Schulz; the optimizer Qwen3.8-Flash-Next / LongCat report) help even in knowledge-limited distillation fine-tunes — unmeasured territory for diff --git a/docs/paper.adoc b/docs/paper.adoc index 0eb1735..f1da9b1 100644 --- a/docs/paper.adoc +++ b/docs/paper.adoc @@ -254,16 +254,18 @@ The engineering consequence: the client tier ships at ByT5-small, quantized. A t A fourth question completes the frontier analysis: at the shipped rung, is the residual teacher–student gap a *parameter-capacity* gap? Recent sparse-memory architectures (product-key memory, Lample et al. 2019; embedding scaling, Liu et al. 2026) argue parameters and compute are separable — lookup capacity at near-zero FLOPs. We tested this directly on the Arabic client tier with a single-variable design: identical teacher, corpus, teacher-labels, seed, and schedule as the shipped student, plus three product-key memory layers (+85.9M parameters, 16,384 slots each, top-32 reads) injected behind zero-initialized gates on the decoder FFNs (the pretrained function is preserved exactly at step 0, and the gates are verified to engage — settling at 0.03–0.05). -[%autowidth,cols="1,1,1"] +[%autowidth,cols="1,1,1,1"] |=== -|Student (Arabic, identical conditions) |Params |DER-CE (full 1,200) +|Student (Arabic, identical teacher/corpus/labels/seed) |Optimizer |Params |DER-CE (full 1,200) -|ByT5-small (shipped) |300M |8.259 -|**+ 3 PKM memory layers** |386M |**7.555** -|r6 teacher |580M |2.582 +|ByT5-small (shipped) |AdamW |300M |8.259 +|ByT5-small |Muon |300M |5.295 +|+ 3 PKM memory layers |AdamW |386M |7.555 +|**+ 3 PKM memory layers** |**Muon** |386M |**4.829** +|r6 teacher |— |580M |2.582 |=== -Memory closes 0.70pp of the 5.68pp teacher–student gap — a real, controlled improvement at near-zero added compute, but below the pre-registered 1.0pp bar under which we would have called the capacity axis *the* fix. The same pre-registration discipline then isolated the *optimization* term: an identical arm differing only in optimizer — Muon (orthogonalized momentum via Newton–Schulz) on the hidden matrices, AdamW for embedding-like parameters — scored **4.829**, a further −2.73pp. The gap at this rung decomposes into ≈0.7pp capacity + ≈2.7pp optimization + ≈2.2pp residual (domain coverage; the full-benchmark subset lesson of <> measures exactly this exposure). Two practical notes travel with the number: Muon's training CE was ~3× lower at equal steps *and* ~2.8× faster per step on this workload (1450-byte windows dominate; Newton–Schulz is cheap next to them), and the adopt gate (≥0.3pp, set before training) was exceeded nine-fold. Optimization — not parameter count — is the dominant recoverable term of the distillation gap at this rung; capacity is a modest second lever, now measured. (This speaks to the *distillation* gap on a pretrained backbone; it does not reopen the from-scratch collapse of finding 1, which is an initialization effect.) +The 2×2 closes cleanly: the optimizer alone recovers −2.96pp, memory alone −0.70pp (−0.47 under Muon), both together −3.43pp — roughly additive, slightly sub-additive on the memory term. The gap at this rung decomposes into ≈3.0pp optimization + ≈0.6pp capacity + ≈2.2pp residual (domain coverage; the full-benchmark subset lesson of <> measures exactly this exposure). Two practical notes travel with the numbers: Muon's training CE was ~3× lower at equal steps *and* ~2.8× faster per step on this workload (1450-byte windows dominate; Newton–Schulz is cheap next to them), and the optimizer adopt gate (≥0.3pp, set before training) was exceeded nine-fold — while the capacity arm, measured under the same discipline, stayed below its 1.0pp bar. Optimization — not parameter count — is the dominant recoverable term of the distillation gap at this rung. (This speaks to the *distillation* gap on a pretrained backbone; it does not reopen the from-scratch collapse of finding 1, which is an initialization effect.) == The decode protocol is part of the measurement [[section-decode]] From c0156bf09d361ae45158c03e558e6c4e953b5e85 Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Fri, 28 Aug 2026 21:58:54 +0800 Subject: [PATCH 4/6] =?UTF-8?q?docs:=20run-005-muon=20recorded=20=E2=80=94?= =?UTF-8?q?=20the=202x2=20factorial=20closes=20at=205.2945?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit --- docs/RESULTS.md | 15 +++++++++++++++ 1 file changed, 15 insertions(+) diff --git a/docs/RESULTS.md b/docs/RESULTS.md index 756b92c..c64dae9 100644 --- a/docs/RESULTS.md +++ b/docs/RESULTS.md @@ -260,6 +260,21 @@ gap at this rung decomposes: ~0.70pp capacity + ~2.73pp optimization + ~2.25pp residual (domain coverage). The vanilla+Muon factorial cell (run-005-muon) completes the decomposition. +### run-005-muon — factorial cell 4: vanilla + Muon (2026-08-28) + +| Model | DER-CE (full 1,200) | +|---|---| +| ByT5-small + Muon (run-005) | 5.2945% | +| ByT5-small + PKM + Muon (run-004) | 4.8287% | +| ByT5-small + PKM + AdamW (run-003) | 7.5553% | +| ByT5-small vanilla + AdamW (run-002) | 8.2590% | + +The 2x2 closes cleanly: optimizer alone −2.96pp; memory alone −0.70pp +(−0.47 under Muon); combined −3.43pp (60.4% of the 5.677pp canonical +gap) — roughly additive, slightly sub-additive on memory. Residual +~2.2pp is domain coverage. Optimization is the dominant recoverable +term of the distillation gap at the ByT5-small rung. + Leaderboard context (SadeedDiac-25, Misraj evaluator, zero-skip, harakat-projected DER-CE): the teacher tier (r6, 580M) at 2.5793 (reproduced at 2.5815, 2026-08-26) is the best dedicated model measured From 05c5fba8fa7797dfaebbdad057863a3f5e0bacd2 Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Fri, 28 Aug 2026 21:59:25 +0800 Subject: [PATCH 5/6] docs: factorial closed in publication notes --- docs/PUBLICATION-NOTES.md | 2 +- 1 file changed, 1 insertion(+), 1 deletion(-) diff --git a/docs/PUBLICATION-NOTES.md b/docs/PUBLICATION-NOTES.md index 6f09bf8..767f6e8 100644 --- a/docs/PUBLICATION-NOTES.md +++ b/docs/PUBLICATION-NOTES.md @@ -81,7 +81,7 @@ ByT5-small rung: ≈0.70pp capacity + 2.73pp optimization + 2.25pp residual (domain coverage). The strongest training-methods result in the paper — first controlled Muon measurement on byte-level seq2seq distillation. Paper: frontier section. EXPERIMENTS.md E3, RESULTS.md -run-004. Factorial cell 4 (vanilla+Muon, run-005) queued. +run-004. **Factorial cell 4 (vanilla+Muon, run-005): 5.2945 — the 2×2 closes cleanly** (optimizer alone −2.96pp, memory alone −0.70pp, combined −3.43pp, roughly additive). Paper carries the full table. ### 10. Arabic news-domain adaptation (r7) — COMPLETE 2026-08-28, NEW CANONICAL TEACHER **ID 2.2864/1.3343 (−0.29pp over r6) AND OOD WikiNews-2024 multi-ref From 6de29a9061ba41299a79ea5bd9843e0043747e57 Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Sat, 29 Aug 2026 06:36:25 +0800 Subject: [PATCH 6/6] fix: orchestrator log() crashes on a fresh arm's missing run dir MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The first watch line for an arm whose training hasn't started opens chain_log.jsonl inside a directory that only distill_sequence creates — FileNotFoundError killed the orchestrator right after run-005's eval (exactly when it reached arm 4), and every relaunch died the same way before logging anything. mkdir(parents=True, exist_ok=True). --- src/gpu/modal_distill.py | 6 +++++- 1 file changed, 5 insertions(+), 1 deletion(-) diff --git a/src/gpu/modal_distill.py b/src/gpu/modal_distill.py index dfccb71..9b649b1 100644 --- a/src/gpu/modal_distill.py +++ b/src/gpu/modal_distill.py @@ -1696,7 +1696,11 @@ def qwen_next_chain() -> dict: ROOT = Path("/checkpoints") def log(run: str, event: str) -> None: - with (ROOT / run / "chain_log.jsonl").open("a", encoding="utf-8") as fh: + # mkdir: a fresh arm's run dir does not exist until its training + # creates it — the first watch line must not crash on that + run_dir = ROOT / run + run_dir.mkdir(parents=True, exist_ok=True) + with (run_dir / "chain_log.jsonl").open("a", encoding="utf-8") as fh: fh.write(json.dumps({"t": round(time.time()), "event": event}) + "\n") CHECKPOINTS.commit()