Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
36 changes: 36 additions & 0 deletions TODO.publish/01-glm-4-7-row-completion.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,36 @@
# 01 — glm-4.7-flash SadeedDiac-25 row completion

Status: IN FLIGHT (2026-09-02). Fetch state: 992/1,200 distinct rows
landed, 208 still empty after three passes — the endpoint sits behind
sustained 429s (code 1305; thinking-disabled accepted, plain
completion inexpressible). Clean single fetch pass running
(/tmp/glm47_eval4.log; the #65 guard drops the 208 empties on
resume). Interim tables from contaminated interleaved passes are
VOID — only the final full-fetch tables count.

## Protocol

- model glm-4.7-flash, effort: none (display label
"thinking-disabled" = the #62 payload path: thinking.type=disabled)
- same 1,200-paragraph SadeedDiac-25 sweep, greedy contract, raw +
projected zero-skip, Misraj evaluator, full disclosure rows

## Remaining steps

- [ ] Fetch to todo=0 (retry passes until provider recovers; each
pass is checkpoint-resumable)
- [ ] Final raw + zero-skip tables (single clean process, no
interleaved writers)
- [ ] results/sadeed-glm-4-7-flash/{README.md, preds CSVs} +
RESULTS.md row + protocol ledger row + paper.adoc row
- [ ] Attribution decomposition (missing/wrong/extra + U+0670 rules;
r7 + GLM-5.2 controls)
- [ ] Paired bootstrap vs GLM-5.2
- [ ] If provider stays blocked after sustained retries: record as
partial-coverage row with explicit disclosure (owner call, see
TODO.publish/05)

## Unblock

Completing this unblocks TODO.publish/04 (site frontier sentence
needs the full regression axis: 5.2 -> 4.7 -> 5.3 -> 5.3-Flash).
30 changes: 30 additions & 0 deletions TODO.publish/02-e6-verdict.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,30 @@
# 02 — E6 verdict (constant-budget register diversification)

Status: IN FLIGHT (2026-09-02). Training COMPLETE — all 11,073 steps,
final CE ~0.003-0.008 (run-008-tashkeela-mix; /tmp/e6_distill2.log).
`modal_distill.py::main` does not chain the final eval, so
`evaluate_der` was launched separately (/tmp/e6_evalder.log,
resumable, writes final_eval.json).

## Registered (EXPERIMENTS.md, gate before launch)

- control: run-006-r7-muon verbatim schedule; delta = 8,000 Tashkeela
classical units replace 8,000 news units (30k total, identical
steps/schedule)
- gate: adopt at <= 4.5218 full-set windowed DER; honest report in
[4.5218, 4.8218); investigate if worse
- prediction: 4.45-4.75
- data hygiene: 120,000 decontaminated units, 48 contaminated
dropped, 17,036 dups removed (rababa-tashkeela v1.0)

## Remaining steps

- [ ] evaluate_der full-set verdict vs gate
- [ ] EXPERIMENTS.md E6 status + RESULTS.md
- [ ] PUBLICATION-NOTES: the E5(-)/E6 rung pair — MTP-aux failed at
+0.26pp (confound disclosed), data-side diversification is the
data-vs-architecture answer either way
- [ ] If adopted: ara-diac-small-2.x material — release is an owner
version decision (see TODO.publish/05)
- [ ] Record training CE curve vs run-006 at matched steps (register
shift visible in convergence speed is itself a finding)
34 changes: 34 additions & 0 deletions TODO.publish/03-d-nikud-vendor-rows.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# 03 — D-Nikud vendor rows (Hebrew Dicta ledger)

Status: COMPLETE (2026-09-02) — with the premise CORRECTED: D-Nikud
publishes NO numbers on the three Dicta ACL 2020 corpora (verified
against arXiv 2402.00075: their tables are internal 5% splits and
the Nakdimon test pipeline; repo README has no numbers). What was
recorded instead: their Nakdimon test-set table (5 systems,
DEC/CHA/WOR/VOC) as vendor-published ledger rows + the cross-metric
non-comparability disclosure + the angle-bracket matres-lectionis
protocol note. See rababa/docs/RESULTS.md Dicta section.

## Design

- D-Nikud (arXiv 2402.00075, NadavShaked/D_Nikud: T5BERT + Bi-LSTM,
~1.5M-token dataset) published numbers on the same three public-
domain Dicta sets (ACL 2020 test corpora) — add as vendor-published
rows beside ours in the RESULTS.md Hebrew Dicta section.
- Vendor rows are context, not protocol-equal comparison: their
metrics are character-level diacritics accuracy (plus nikud-letter
handling per their paper), not our DER; their decode is theirs.
The ledger row must disclose this asymmetry explicitly (the
standard since the GLM-5.3 attribution correction: never let a
cross-protocol number stand next to ours unqualified).
- License: test corpora public domain; D-Nikud numbers cited from
the paper with arXiv ref.

## Remaining steps

- [ ] Pull D-Nikud's paper table numbers for the three Dicta sets
(verify against the paper, not blog/README summaries)
- [ ] Vendor rows + protocol-disclosure note in rababa/docs/
RESULTS.md Hebrew section
- [ ] Cross-ref in paper.adoc Hebrew section if the numbers appear
there too
28 changes: 28 additions & 0 deletions TODO.publish/04-site-frontier-sentence.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,28 @@
# 04 — Site frontier sentence (interscript.org ml page)

Status: BLOCKED on TODO.publish/01 (needs the full regression axis).
The ml.astro page exists on interscript.org (rebrand 2026-08); what
it lacks is the one-line frontier-context claim that the GLM sweep
now substantiates.

## The claim (draft, number slots to fill from 01)

Frontier LLMs regress on classical-knowledge diacritization as their
optimization shifts agentic: on SadeedDiac-25, GLM-5.2 {5.2-DER} ->
glm-4.7-flash {4.7-DER} -> GLM-5.3 {9.90} -> GLM-5.3-Flash {8.80}
raw, versus dedicated distilled students at {r6-DER} — the dedicated
model is not nostalgia, it is the only thing on the right side of
the axis. Attribution: the regression is wrong-haraqat, not
convention (U+0670 = 0.125pp total; rababa #69 / ml #116).

- One sentence + the numbers, linking to the protocol ledger
(rababa docs/RESULTS.md) for the disclosed rows
- Bootstrap CIs quoted from the ledger (paired vs 5.2)

## Remaining steps

- [ ] Fill slots from 01's final tables (or record partial-coverage
disclosure if that is the outcome)
- [ ] PR on interscript.org release-branch pattern (never main)
- [ ] Mirror check: the sentence must match ledger numbers exactly —
never-fabricate rule applies to marketing copy too
34 changes: 34 additions & 0 deletions TODO.publish/05-owner-decisions.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,34 @@
# 05 — Owner-decision register (publish campaign)

Status: REGISTERED (2026-09-02). Everything in this campaign that is
the owner's call, surfaced in one place. Cross-reference: the fuller
register with rationale lives in TODO.training-work/08; this file is
the publish-facing view. Nothing here is executed without an explicit
instruction naming it.

## Decisions pending

1. **GKD ordering** — the last registered training lever (SOTA
strategy 2025: "GKD is next"). Runs after E6's verdict so the
rung ladder stays one-variable-at-a-time. Owner decides: launch
now, after E6, or wait for the E5/E6 writeup.
2. **glm-4.7 partial-coverage policy** — if the 429 wall never
clears: record as partial-coverage row with disclosure, or drop
the row. (Default per protocol honesty: disclose coverage in the
row itself.)
3. **ara-diac-small-2.x release** — iff E6 passes the gate
(<= 4.5218). Version number is always the owner's decision
(rubygems lesson applies to model indices too).
4. **head32 swap-in shape** — in-place + index-v2 vs parallel
`-int8-head32` ids. Five-of-five rebuilds done with flip CIs;
swap-in is a release-side act.
5. **fp16 index entries for ara-diac-2.0** — export-side, low risk,
owner sequencing.
6. **rababa PR backlog** — #51/#52/#53/#48/#49/#26 (from
TODO.training-work/08).

## Standing constraints (do not re-derive)

- Version numbers, releases, tags: owner only.
- PRs rebase-merge only on explicit instruction; never main.
- Protocol-matched numbers only; partial coverage disclosed in-row.
68 changes: 68 additions & 0 deletions TODO.qwen-next/01-margin-aware-parity-gates.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,68 @@
# 01 — Margin-aware parity gates (adopt now)

Source: Qwen3.8-Flash-Next quantization practice (Unsloth day-0 analysis) —
the dense backbone tolerates aggressive quantization, but random-access
tensors (embedding tables, output heads, PLE lookup tables) degrade fast
below 4-bit. Precision floor is a property of **how a tensor is read**,
not just its size. They validate with KL-divergence on logits, not
output equality alone.

## Gap in our stack

`ml-models/src/imf/parity.py` gates releases on CER-delta between the
torch reference and the ONNX zip (0.2pp fp32 → 3.0pp int4). Our byte
students have **flat top-1 margins** (the greedy-vs-beam lesson: tha-small
published 12.06 PER was really 2.85). On flat distributions, quantization
noise below exact-match detection can still flip near-tie argmaxes — a
golden-set CER gate can pass while decode fragility ships.

## Deliverables

1. `run_margin_analysis(model, zip_path, pairs)` in `src/imf/parity.py`:
teacher-forced forward on both sides (torch decoder vs ONNX decoder
session), per-token top1−top2 logit margins, argmax flip rate, KLD.
2. `MarginReport` dataclass + JSON serialization; written next to the zip
as `<id>-margins-<precision>.json` (diagnostic, not a release blocker
yet — schema churn avoided).
3. Unit test on the fixture checkpoint (`make_fixture_checkpoint`).
4. Real validation: khm-latn-1.0-fp16.zip (local) + one int8/int4 zip.
5. Policy (documented in docs/EXPERIMENTS.md): embeddings + lm_head +
any memory tables are a separate quantization class — stay fp16 (≥int8
at minimum) when the body is quantized.

## Protocol (paper-ready)

- Probe set: the parity golden inputs (same pairs the CER gate uses).
- Metrics: flip_rate = fraction of teacher-forced positions where
argmax(torch) ≠ argmax(onnx); margin quantiles (p1/p10/p50) of the
reference; KLD mean; per-precision.
- Hypothesis: flip_rate correlates with cer_delta but detects fragility
earlier; int4 artifacts show near-tie flips invisible to the CER gate.

## Status

- [x] run_margin_analysis + MarginReport in parity.py
- [x] Wired into modal_export parity gate (every export emits margins)
+ read-only `margins` entrypoint for published zips
- [x] Unit test (fixture checkpoint, synthetic near-tie logits) — 8 pass
- [x] **Full-catalog table measured (12 rows, docs/EXPERIMENTS.md E1)**:
fp32 exact everywhere; fp16 benign except khm 2.93% (flattest
margins, p50 0.121); int8 0.26–9.34%. **heb-diac-1.1 int8 = the
outlier: 9.34% flips, only 20% near-tie** — 80% of argmax flips at
confident positions, invisible to the CER gate. Open item:
re-examine the Hebrew int8 artifact (per-channel quantization or
serve fp16) + add a margin threshold to the release policy.
tha-g2p-small int4: 0.26%, all near-tie — flat-but-consistent.
- [x] **Root cause found + export default fixed (fcbe5d3)**: the outlier
was the quantized *head*. Probes on the same pairs: per-channel
alone 8.50% (+25% size, rejected); **head-fp32 body-int8 → 0.26%
flips (36×), KLD 47× lower, 100% near-tie, +0.4% size**.
`export_zips` now excludes `/lm_head/MatMul` from int8 by default.
Shipped int8 zips predate it; re-export = release decision (open).
- [x] Policy registered in docs/EXPERIMENTS.md (E1)
- [x] **Corrected-artifact sweep (2026-08-29, authorized)**:
`rebuild_int8_head32` (PR #76) rebuilds every shipped int8 zip
with the head in fp32 and gates it (parity in-zip + margin
budget). Artifacts land as {mid}-int8-head32.zip; swap-in is a
version decision pending results. Sweep running across khm/urd×2/
heb/tha; tha int4 left as-is (benign: 0.26%, all near-tie).
74 changes: 74 additions & 0 deletions TODO.qwen-next/02-pkm-memory-layer-student.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,74 @@
# 02 — PKM memory-layer student (client-tier v2, the big one)

Sources:
- arXiv 2601.21204 "Scaling Embeddings Outperforms Scaling Experts"
(LongCat-Flash-Lite: 68.5B params, >30B in embedding/lookup tables,
~3B active — embedding scaling beats expert scaling on the Pareto
frontier in specific regimes).
- Qwen3.8-Flash-Next: first public model on that recipe (51B N-gram/PLE
tables, ~6B active).
- Precedent at our modality: product-key memory was demonstrated on
**character-level** LM (Lample et al., NeurIPS 2019, enwik8/text8).

## Why this is the right experiment for us

Wrong turn 1 measured: sub-100M from-scratch byte students collapse; a
pretrained backbone is non-negotiable; the client tier is pinned at
ByT5-small 300M → 246 MiB int8 (Arabic student: 8.26 DER full-set vs
teacher 2.58 — a 5.68pp gap).

The LongCat result reframes the axis: **parameters and compute are
separable**. Keep compute at ByT5-small, add lookup capacity. Haraqat
restoration is heavily lexical knowledge (word identity + local context)
— table-friendly, not compute-friendly. It is also the architectural
version of our r6/r8 aux-task finding: tagged knowledge injection works;
a memory layer is where that knowledge can physically live.

## Design (as launched)

- Backbone: `google/byt5-small` (pretrained, per the non-negotiable rule).
Measured at launch: 300M base, d_model 1472, **4 decoder blocks**
(ByT5 depth lives in the encoder) — memory layers on decoder blocks
[-1, -2, -3].
- Memory: product-key factorization, C=128 keys/half → 16,384 slots,
top-k=32 sparse reads → **+85.9M params (+29%)** at near-zero FLOPs.
Zero-init output gate preserves the pretrained function at step 0
(unit-tested: logits bit-identical pre/post injection).
- Training: identical to ara-diac-small run-002 — teacher r6 frozen,
same r5-units corpus (29,322 pairs), same `teacher_labels_v2.jsonl`
(copied into the run dir; labels trusted complete — no relabeling),
same seed → **single-variable comparison**.

## Protocol + gate (pre-agreed)

- Spec `ara-diac-small-pkm` in ml-models/src/gpu/modal_distill.py.
- Gate: windowed zero-skip Misraj DER-CE, full 1,200 paragraphs,
≤ 3.07 (the run-002 gate) — and the comparison target is run-002's
8.259 full-set student / 2.5815 teacher.
- Verdict rule: PKM wins if it closes ≥1.0pp of the 5.68pp gap at equal
decode-time compute (memory reads are gathers, not matmuls). If no
movement, the capacity story is wrong — the gap is modeling/optimization,
publishable as a negative either way.
- Onnx export (opset-14 TopK/Gather) only if it wins; noted, not built yet.

## Launch

- A10G distill slot (never competes with A100 teacher runs) — fits under
the ≤2-big-GPU-apps budget alongside r8. 10,995 steps, ~7h.
- Detached; chain watchdog crash-relaunches, then runs
evaluate_der, then launches the Muon A/B arm (TODO 03), then evals it.

## Status

- [x] pkm.py (PKMLayer + byt5 injection helper) + smoke tests (6 pass)
- [x] Spec wiring in modal_distill.py (train + eval load paths)
- [x] Launched detached: rababa_arabic_distill_small/run-003-pkm
- [x] Registered in ml-models docs/EXPERIMENTS.md (E2)
- [x] Engagement probe (`pkm_gates`): step-500 gates 0.0008/0.0028/0.0019
— all three off zero; memory branch engaged (CE 2.07 → 0.086 @ 900)
- [x] **VERDICT (2026-08-28): 7.5553 full-set DER vs run-002's 8.259 —
0.704pp of the 5.677pp gap closed (12.4% relative), below the
pre-registered ≥1.0pp win bar.** Positive direction, honestly
reported; capacity is real but not the dominant term. Gate
engagement verified (gates off zero, CE 2.07→0.02). Muon arm (E3)
now tests the optimization half of the remaining gap.
46 changes: 46 additions & 0 deletions TODO.qwen-next/03-muon-optimizer-ab.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,46 @@
# 03 — Muon optimizer A/B (one cheap test, low expected gain)

Source: Qwen3.8-Flash-Next / LongCat-Flash-Lite trained with Muon
(orthogonalized momentum via Newton–Schulz); "~1/9th training cost" claims
circulate for from-scratch pretraining.

Our prior: fine-tunes are **knowledge-limited, not optimization-limited**
(RL was a measured negative; data/supervision-side levers won). So the
expected gain on a 3-epoch distill is small — but the test is cheap and
runs in the same A10G slot after the PKM run.

## Design

- `muon.py` in ml-models/src/gpu/: single-file Muon — Newton–Schulz
orthogonalization applied to 2D weight matrices (hidden/FFN/attn
projections); embeddings, layer norms, and the lm_head stay on AdamW
(standard split; embeddings are not orthogonalizable objects).
- Flag `optimizer: muon|adamw` in the distill spec (default adamw —
nothing changes for existing specs).
- A/B: same spec as TODO 02 (`ara-diac-small-pkm`), same seed 42, only
the optimizer differs → run-003-pkm (adamw) vs run-004-pkm-muon.

## Protocol + gate

- Metric: windowed DER-CE on the val slice during training + full
1,200-paragraph Misraj at the end; wall-clock per step recorded.
- Adopt if: ≥0.3pp DER improvement at equal steps AND no stability
regressions (loss spikes, grad-norm blowups) AND step overhead <15%.
- Either direction is paper-reportable (training-methods appendix):
"Muon vs AdamW on byte-level distillation fine-tunes" is unmeasured
territory for seq2seq students.

## Status

- [x] muon.py (Newton–Schulz, param-group split; shared.weight routing
fixed for transformers 5.x) + unit tests
- [x] Spec `ara-diac-small-pkm-muon` (run-004) wired in modal_distill.py
- [x] **VERDICT (2026-08-28): 4.8287 vs 7.5553 — −2.727pp from the
optimizer alone; adopt gate (≥0.3pp) exceeded 9×. ADOPTED.**
Training CE ~0.007 vs ~0.02 at equal steps; ~1.2s/step vs ~3.4s;
no stability events. Gap decomposition: ≈0.70pp capacity +
2.73pp optimization + 2.25pp residual.
- [x] Factorial cell 4 (vanilla+Muon, run-005-muon) landed 2026-08-28:
**5.2945** — the 2×2 closes cleanly (optimizer alone −2.96pp,
memory alone −0.70pp / −0.47 under Muon, combined −3.43pp).
Paper carries the full factorial table.
22 changes: 22 additions & 0 deletions TODO.qwen-next/04-speculative-decoding-probe.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,22 @@
# 04 — Speculative decoding probe (PARKED)

Source: LongCat-Flash-Lite converts embedding sparsity into inference
speed via speculative decoding; Qwen ships day-0 vLLM support.

Ours: byte-level outputs are long (1400B window → up to 3200 tokens), so
the theory applies — a tiny byte draft model + the student as verifier
could cut server-tier batch decode latency.

Why parked:
- Teachers run offline; the API serves greedy small students that are
already fast for their workloads.
- We have no latency-budget data showing decode time binds at the API.
- The draft model would itself need training + a parity story — not free.

Revisit trigger: API latency metrics showing p95 decode > budget, or a
client-tier memory-layer win (TODO 02) that inflates compute enough for
draft/verify to matter. Registered here so the idea isn't lost.

## Status

- [x] Parked with explicit revisit trigger (no code, by design)
Loading