Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
35 changes: 29 additions & 6 deletions docs/EXPERIMENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -123,12 +123,19 @@ All rows passed the CER parity gate at release. Readings:
equal steps; step time ~1.2s vs ~3.4s (the <15% overhead gate met
with margin — Newton–Schulz is cheap next to 1450-byte windows); no
stability events.
- **Combined with E2, the gap decomposition at the ByT5-small rung:**
capacity ~0.70pp (PKM) + optimization ~2.73pp (Muon, on the PKM
architecture) + residual ~2.25pp (domain coverage — the subset
lesson). Caveat carried: the A/B isolates the optimizer ON the
memory architecture; the vanilla+Muon cell (run-005-muon) is queued
to complete the 2x2 factorial.
- **Factorial cell 4 (run-005-muon, vanilla + Muon, 2026-08-28):
5.2945.** The 2x2 closes cleanly:

| Full-set DER | AdamW | Muon |
|---|---|---|
| vanilla ByT5-small | 8.259 | 5.295 |
| + PKM memory | 7.555 | **4.829** |

Decomposition (teacher ~2.6): optimizer alone −2.96pp; memory alone
−0.70pp (−0.47 under Muon); combined −3.43pp — roughly additive,
slightly sub-additive on the memory term. Optimization is the
dominant recoverable term at this rung; capacity is second; residual
~2.2pp is domain coverage (the subset lesson).
- **Hypothesis:** orthogonalized-momentum updates (Newton–Schulz; the
optimizer Qwen3.8-Flash-Next / LongCat report) help even in
knowledge-limited distillation fine-tunes — unmeasured territory for
Expand All @@ -143,6 +150,22 @@ All rows passed the CER parity gate at release. Readings:
- **Report:** either direction goes to the paper's training-methods
appendix.

## E4 — ara-diac-small-2.0 candidate (run-006-r7-muon)

- **Status:** queued (registered 2026-08-28 before launch).
- **Hypothesis:** the two measured wins compound — the r7 canonical
teacher (better labels; ID 2.2864 vs 2.5793) plus the E3-adopted
Muon optimizer (−2.727pp on r6 labels) — moving the client rung far
below the shipped 8.259.
- **Design:** identical corpus/limits/schedule to run-002/003/004/005;
teacher = run-007-news/best (fresh greedy labels, teacher_labels_r7);
optimizer = Muon (embedding-like group on AdamW); vanilla ByT5-small
(the shipped architecture — the 2.0 ships what was measured).
- **Pre-agreed gate:** beats the shipped student by ≥2.0pp full-set
windowed DER (i.e. ≤ 6.26) to justify the 2.0 release; the strict
teacher+0.5pp budget remains the disclosed north star.
- **Prediction (registered):** 4.3–5.0, by E3's 4.829 on weaker labels.

## Parked

- **Speculative decoding** (LongCat converts sparsity→speed): revisit
Expand Down
20 changes: 11 additions & 9 deletions docs/PUBLICATION-NOTES.md
Original file line number Diff line number Diff line change
Expand Up @@ -81,16 +81,18 @@ ByT5-small rung: ≈0.70pp capacity + 2.73pp optimization + 2.25pp
residual (domain coverage). The strongest training-methods result in
the paper — first controlled Muon measurement on byte-level seq2seq
distillation. Paper: frontier section. EXPERIMENTS.md E3, RESULTS.md
run-004. Factorial cell 4 (vanilla+Muon, run-005) queued.
run-004. **Factorial cell 4 (vanilla+Muon, run-005): 5.2945 — the 2×2 closes cleanly** (optimizer alone −2.96pp, memory alone −0.70pp, combined −3.43pp, roughly additive). Paper carries the full table.

### 10. Arabic news-domain adaptation (r7) — ID LANDED 2026-08-28
**2.2864 / 1.3343 full-set windowed zero-skip — new best dedicated
model, −0.29pp over r6** (the news mix improved in-domain, not just
OOD; behind only Claude-3.7-Sonnet's published 1.3941). OOD half
(WikiNews-2024 multi-ref, gate: beat r6's 19.82/12.46) running via the
auto-launched actor. → rababa RESULTS.md (recorded) → paper leaderboard
table + OOD note once OOD confirms; canonical-teacher promotion
pending that check.
### 10. Arabic news-domain adaptation (r7) — COMPLETE 2026-08-28, NEW CANONICAL TEACHER
**ID 2.2864/1.3343 (−0.29pp over r6) AND OOD WikiNews-2024 multi-ref
17.38/11.83 (vs r6's 19.82/12.46)** — r7 sweeps both surfaces; the
teacher-lineage domain trade-off is gone. Best dedicated model under
the protocol, behind only Claude-3.7-Sonnet's published 1.3941, now
well clear of GLM-5.2 (2.6911). r7 replaces r6 as the canonical
Arabic teacher for future distillations. Paper leaderboard table
updated; rababa RESULTS.md carries both tables. Follow-on (user
decision): ara-diac-2.0 teacher release; students re-distilled from
r7 with Muon (E3) are the natural ara-diac-small-2.0.

## Venue fit (working notes)

Expand Down
15 changes: 15 additions & 0 deletions docs/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -260,6 +260,21 @@ gap at this rung decomposes: ~0.70pp capacity + ~2.73pp optimization
+ ~2.25pp residual (domain coverage). The vanilla+Muon factorial cell
(run-005-muon) completes the decomposition.

### run-005-muon — factorial cell 4: vanilla + Muon (2026-08-28)

| Model | DER-CE (full 1,200) |
|---|---|
| ByT5-small + Muon (run-005) | 5.2945% |
| ByT5-small + PKM + Muon (run-004) | 4.8287% |
| ByT5-small + PKM + AdamW (run-003) | 7.5553% |
| ByT5-small vanilla + AdamW (run-002) | 8.2590% |

The 2x2 closes cleanly: optimizer alone −2.96pp; memory alone −0.70pp
(−0.47 under Muon); combined −3.43pp (60.4% of the 5.677pp canonical
gap) — roughly additive, slightly sub-additive on memory. Residual
~2.2pp is domain coverage. Optimization is the dominant recoverable
term of the distillation gap at the ByT5-small rung.

Leaderboard context (SadeedDiac-25, Misraj evaluator, zero-skip,
harakat-projected DER-CE): the teacher tier (r6, 580M) at 2.5793
(reproduced at 2.5815, 2026-08-26) is the best dedicated model measured
Expand Down
21 changes: 12 additions & 9 deletions docs/paper.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -203,21 +203,22 @@ External claims are made only where a public benchmark exists and the full proto
|System |Params |DER-CE

|Claude 3.7 Sonnet (published) |— |1.3941
|**our teacher (r6, ByT5-base)** |580M |**2.5793**
|**our teacher (r7, ByT5-base)** |580M |**2.2864**
|our previous teacher (r6) |580M |2.5793
|GLM-5.2 (reproduced, zero-skip) |— |2.6911
|Gemini Flash 2.0 (published) |— |3.1926
|GPT-4 (published) |— |3.8645
|Sadeed (published) |1.5B |7.2915
|our client student (ByT5-small) |300M |8.259
|===

The teacher is the best dedicated (task-trained, runnable-locally) model measured under this protocol — second only to a frontier proprietary LLM's published number, ahead of a clean GLM-5.2 reproduction, Gemini Flash, and GPT-4 — and it does so at 580M parameters against Sadeed's 1.5B.
The teacher is the best dedicated (task-trained, runnable-locally) model measured under this protocol — second only to a frontier proprietary LLM's published number, ahead of a clean GLM-5.2 reproduction, Gemini Flash, and GPT-4 — and it does so at 580M parameters against Sadeed's 1.5B. The current rung (r7) adds a teacher-labeled news-domain mix to the morphological-auxiliary teacher (r6) and improves both surfaces at once: in-domain 2.5793 → 2.2864, and out-of-domain WikiNews-2024 multi-reference 19.82/12.46 → 17.38/11.83 (WER/DER) — the paragraph-specialization trade-off of the earlier lineage is gone.

The r6 teacher's auxiliary-task design is validated by a controlled ablation: an identical run differing only in the auxiliary stream's output representation — broad-phonemic IPA of the same training units (deterministic converter) in place of morphological analysis — scores 2.6588 (vs 2.6775 for no auxiliary task and 2.5793 for the morphological one), while a held-out probe confirms the phonemic projection itself was learned (2.3% CER on the IPA stream). The improvement is therefore attributable to lexical-morphological knowledge injection, not to phonemic supervision per se: the strong form of the "phonemes help diacritization" hypothesis fails where its weak form (any structured auxiliary projection beats none) barely holds.

The client student's row carries a measurement lesson we record rather than bury. Its first-published figure, 3.66, was measured on the first 300 paragraphs of the benchmark; the full 1,200-paragraph run scores 8.26 (teacher reproduces at 2.5815 against its documented 2.5793, confirming protocol consistency). The subset was not representative: its paragraphs sit closer to the student's training domain, and the remaining 900 expose a domain-generalization gap the subset hid. Two practices follow, now standing rules: student-tier numbers are published only from full benchmark sets, and a subset figure is labeled as such at first publication. The student ships with the full-set number disclosed — behind Sadeed-1.5B on the leaderboard — and closing its domain gap (more diverse label coverage, or on-policy distillation; <<section-discussion>>) is open work.

Two protocol notes the benchmark's own authors make and we adopt: self-reported numbers from other protocols (Sadeed's repository reports 1.2% DER under its own split) are not comparable, and multi-reference evaluation (Mohamed & Mubarak 2025) is the honest treatment of valid-alternative diacritizations; on WikiNews-2024 our teacher family scores 19.82 WER against QCRI's in-domain-trained 2.70 — a domain gap (classical hadith vs news text), recorded rather than averaged away.
Two protocol notes the benchmark's own authors make and we adopt: self-reported numbers from other protocols (Sadeed's repository reports 1.2% DER under its own split) are not comparable, and multi-reference evaluation (Mohamed & Mubarak 2025) is the honest treatment of valid-alternative diacritizations; on WikiNews-2024 our teacher scores 17.38 WER against QCRI's in-domain-trained 2.70 — a domain gap (classical hadith vs news text), recorded rather than averaged away.

*Hebrew.* No external leaderboard exists for the Nakdimon test domain (Biblical/Rabbinic). On that split our best teacher scores 16.58% DER — against DictaBERT-large, the reference production diacritizer, at 35.63% *on the same test, run by us*. Recent open systems (D-Nikud; visual-representation approaches, Elboher & Pinter 2025) report state of the art on *modern*-text benchmarks we do not evaluate. Hebrew remains the campaign's disclosed weak spot in absolute terms while being locally competitive; closing the domain gap is active work.

Expand Down Expand Up @@ -253,16 +254,18 @@ The engineering consequence: the client tier ships at ByT5-small, quantized. A t

A fourth question completes the frontier analysis: at the shipped rung, is the residual teacher–student gap a *parameter-capacity* gap? Recent sparse-memory architectures (product-key memory, Lample et al. 2019; embedding scaling, Liu et al. 2026) argue parameters and compute are separable — lookup capacity at near-zero FLOPs. We tested this directly on the Arabic client tier with a single-variable design: identical teacher, corpus, teacher-labels, seed, and schedule as the shipped student, plus three product-key memory layers (+85.9M parameters, 16,384 slots each, top-32 reads) injected behind zero-initialized gates on the decoder FFNs (the pretrained function is preserved exactly at step 0, and the gates are verified to engage — settling at 0.03–0.05).

[%autowidth,cols="1,1,1"]
[%autowidth,cols="1,1,1,1"]
|===
|Student (Arabic, identical conditions) |Params |DER-CE (full 1,200)
|Student (Arabic, identical teacher/corpus/labels/seed) |Optimizer |Params |DER-CE (full 1,200)

|ByT5-small (shipped) |300M |8.259
|**+ 3 PKM memory layers** |386M |**7.555**
|r6 teacher |580M |2.582
|ByT5-small (shipped) |AdamW |300M |8.259
|ByT5-small |Muon |300M |5.295
|+ 3 PKM memory layers |AdamW |386M |7.555
|**+ 3 PKM memory layers** |**Muon** |386M |**4.829**
|r6 teacher |— |580M |2.582
|===

Memory closes 0.70pp of the 5.68pp teacher–student gap — a real, controlled improvement at near-zero added compute, but below the pre-registered 1.0pp bar under which we would have called the capacity axis *the* fix. The same pre-registration discipline then isolated the *optimization* term: an identical arm differing only in optimizer — Muon (orthogonalized momentum via Newton–Schulz) on the hidden matrices, AdamW for embedding-like parameters — scored **4.829**, a further −2.73pp. The gap at this rung decomposes into ≈0.7pp capacity + ≈2.7pp optimization + ≈2.2pp residual (domain coverage; the full-benchmark subset lesson of <<section-leaderboard>> measures exactly this exposure). Two practical notes travel with the number: Muon's training CE was ~3× lower at equal steps *and* ~2.8× faster per step on this workload (1450-byte windows dominate; Newton–Schulz is cheap next to them), and the adopt gate (≥0.3pp, set before training) was exceeded nine-fold. Optimization — not parameter count — is the dominant recoverable term of the distillation gap at this rung; capacity is a modest second lever, now measured. (This speaks to the *distillation* gap on a pretrained backbone; it does not reopen the from-scratch collapse of finding 1, which is an initialization effect.)
The 2×2 closes cleanly: the optimizer alone recovers −2.96pp, memory alone −0.70pp (−0.47 under Muon), both together −3.43pp — roughly additive, slightly sub-additive on the memory term. The gap at this rung decomposes into ≈3.0pp optimization + ≈0.6pp capacity + ≈2.2pp residual (domain coverage; the full-benchmark subset lesson of <<section-leaderboard>> measures exactly this exposure). Two practical notes travel with the numbers: Muon's training CE was ~3× lower at equal steps *and* ~2.8× faster per step on this workload (1450-byte windows dominate; Newton–Schulz is cheap next to them), and the optimizer adopt gate (≥0.3pp, set before training) was exceeded nine-fold — while the capacity arm, measured under the same discipline, stayed below its 1.0pp bar. Optimization — not parameter count — is the dominant recoverable term of the distillation gap at this rung. (This speaks to the *distillation* gap on a pretrained backbone; it does not reopen the from-scratch collapse of finding 1, which is an initialization effect.)

== The decode protocol is part of the measurement
[[section-decode]]
Expand Down
39 changes: 39 additions & 0 deletions models/ara-diac/ara-diac-2.0.README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,39 @@
# ara-diac-2.0

Arabic diacritization (haraqat restoration), ByT5-base, 580M parameters.

## What changed from 1.0

ara-diac-1.0 (r6) was trained with a morphological auxiliary task
(qalsadi iʿrāb labels) on top of the paragraph-context r5 lineage.
ara-diac-2.0 (r7) adds a news-domain adaptation stage: 13,986
teacher-labeled news units mixed at 0.85% with the r5-units anchor,
plus 400 gold WikiNews-2014 lines, initialized from r6.

The result is a full sweep over 1.0 — both surfaces improved:

| Surface | ara-diac-1.0 (r6) | ara-diac-2.0 (r7) |
|---|---|---|
| SadeedDiac-25, full 1,200, windowed zero-skip, greedy | DER 2.5793 / Morph 1.5317 | **DER 2.2864 / Morph 1.3343** |
| WikiNews-2024 multi-ref (QCRI protocol, full) | WER 19.8191 / DER 12.4613 | **WER 17.3794 / DER 11.8273** |

The earlier lineage's in-domain/out-of-domain trade-off is gone: the
news mix improved in-domain substantially, not just OOD.

## Positioning

Best dedicated (task-trained, runnable-locally) model measured on
SadeedDiac-25 under the zero-skip Misraj protocol — behind only
Claude-3.7-Sonnet's published 1.3941, ahead of GLM-5.2 (2.6911
zero-skip reproduction), Gemini-Flash-2.0 (3.1926), and GPT-4 (3.8645),
at 580M parameters against Sadeed-1.5B's 7.2915.

## Contract

IMF v1 zip: byte tokenizer (id = byte + 3, trailing EOS), KV-cache
decoder, opset 14. Greedy decode is the runtime protocol. int8 zips
keep the tied head in fp32 (the margin-gate fix); every precision is
gated by CER parity AND the argmax-flip margin budget.

Provenance: rababa `train_arabic_r7.py`, run `rababa_arabic_byt5/run-007-news`;
verdict tables in `rababa/docs/RESULTS.md` (r7 sections). BSD-3-Clause.
33 changes: 33 additions & 0 deletions models/ara-diac/ara-diac-2.0.metadata.yaml
Original file line number Diff line number Diff line change
@@ -0,0 +1,33 @@
format: imf-v1
id: ara-diac-2.0
task: diacritization
source_script: Arab
target: Arab
tokenizer: bytes
opset: 14
decoder: kv
precision: fp32
license: BSD-3-Clause
trained_from: 'rababa train_arabic_r7.py run-007-news (news-domain adaptation of the r6
morphological-aux teacher: r5-units anchor + 13,986 teacher-labeled news units at
0.85% mix + 400 gold WikiNews-2014 lines, init from run-006-morph); checkpoint
rababa-checkpoints:/rababa_arabic_byt5/run-007-news/best'
metrics:
- name: der_total_greedy
value: 2.2864
protocol: greedy decode (the v1 runtime path); SadeedDiac-25, windowed zero-skip
at 1400 bytes, full 1,200 paragraphs; ByT5-base r7
source: rababa/docs/RESULTS.md#r7-verdict-table-sadeeddiac-25-2026-08-28
- name: der_morph_greedy
value: 1.3343
protocol: same harness; Morphological DER (word-final case endings only)
source: rababa/docs/RESULTS.md#r7-verdict-table-sadeeddiac-25-2026-08-28
- name: wer_wikinews_multiref
value: 17.3794
protocol: WikiNews-2024 multi-reference, QCRI EvalDiac protocol (full mode), 356
texts / 10,616 words
source: rababa/docs/RESULTS.md#r7-ood-verdict-table-wikinews-2024-multiref-2026-08-28
- name: der_wikinews_multiref
value: 11.8273
protocol: same harness
source: rababa/docs/RESULTS.md#r7-ood-verdict-table-wikinews-2024-multiref-2026-08-28
26 changes: 25 additions & 1 deletion src/gpu/modal_distill.py
Original file line number Diff line number Diff line change
Expand Up @@ -205,6 +205,25 @@ def _ensure_src_path() -> None:
"mode": "sequence",
"note": "vanilla ByT5-small + Muon (factorial cell 4)",
},
"ara-diac-small-2": {
# ara-diac-small-2.0 candidate: the r7 canonical teacher + the
# E3-adopted Muon optimizer. Fresh labels from r7 (the v2 labels
# are r6-teacher). EXPERIMENTS.md E4.
"teacher": "rababa_arabic_byt5/run-007-news/best",
"teacher_volume": "rababa",
"student_init": "google/byt5-small",
"optimizer": "muon",
"muon_lr": "0.01",
"train": "r5-units/domain.txt",
"train_extra": ["r5-units/replay.txt"],
"unit_limits": [24000, 6000],
"max_len": 1450,
"label_beams": "1",
"out": "rababa_arabic_distill_small/run-006-r7-muon",
"labels_file": "teacher_labels_r7.jsonl",
"mode": "sequence",
"note": "r7 teacher + Muon; E4 gate: beat the shipped 8.259 by >= 2pp",
},
"ara-diac-tiny": {
"teacher": "rababa_arabic_byt5/run-006-morph/best",
"teacher_volume": "rababa",
Expand Down Expand Up @@ -1672,11 +1691,16 @@ def qwen_next_chain() -> dict:
("ara-diac-small-pkm", "rababa_arabic_distill_small/run-003-pkm"),
("ara-diac-small-pkm-muon", "rababa_arabic_distill_small/run-004-pkm-muon"),
("ara-diac-small-muon", "rababa_arabic_distill_small/run-005-muon"),
("ara-diac-small-2", "rababa_arabic_distill_small/run-006-r7-muon"),
]
ROOT = Path("/checkpoints")

def log(run: str, event: str) -> None:
with (ROOT / run / "chain_log.jsonl").open("a", encoding="utf-8") as fh:
# mkdir: a fresh arm's run dir does not exist until its training
# creates it — the first watch line must not crash on that
run_dir = ROOT / run
run_dir.mkdir(parents=True, exist_ok=True)
with (run_dir / "chain_log.jsonl").open("a", encoding="utf-8") as fh:
fh.write(json.dumps({"t": round(time.time()), "event": event}) + "\n")
CHECKPOINTS.commit()

Expand Down
9 changes: 9 additions & 0 deletions src/gpu/modal_export.py
Original file line number Diff line number Diff line change
Expand Up @@ -87,6 +87,15 @@
"test_data": "nakdimon/test-imf.jsonl",
"probe": "שלום",
},
"ara-diac2": {
"volume": "/volumes/rababa-checkpoints",
"checkpoint": "rababa_arabic_byt5/run-007-news/best",
"metadata": "models/ara-diac/ara-diac-2.0.metadata.yaml",
"readme": "models/ara-diac/ara-diac-2.0.README.md",
"test_volume": "/datasets/rababa",
"test_data": "arabic-sadeed-imf/test.jsonl",
"probe": "مكتبة",
},
"ara-diac": {
"volume": "/volumes/rababa-checkpoints",
"checkpoint": "rababa_arabic_byt5/run-006-morph/best",
Expand Down
Loading