Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
20 changes: 13 additions & 7 deletions TODO.publish/01-glm-4-7-row-completion.md
Original file line number Diff line number Diff line change
@@ -1,12 +1,18 @@
# 01 — glm-4.7-flash SadeedDiac-25 row completion

Status: IN FLIGHT (2026-09-02). Fetch state: 992/1,200 distinct rows
landed, 208 still empty after three passes — the endpoint sits behind
sustained 429s (code 1305; thinking-disabled accepted, plain
completion inexpressible). Clean single fetch pass running
(/tmp/glm47_eval4.log; the #65 guard drops the 208 empties on
resume). Interim tables from contaminated interleaved passes are
VOID — only the final full-fetch tables count.
Status: COMPLETE (2026-09-02). 1,200/1,200, zero empties after 12
resumable passes past sustained 429s (sentinels 249 -> 208 -> 79 ->
37 -> 20 -> 13 -> 6 -> 4 -> 2 -> 0). Final: **13.0035 raw /
13.2256 zero-skip** (w/o-CE 10.0510/10.3206; NFDW 17.60 raw).
Attribution (validated scorer — reproduces the #69 series exactly:
5.2 wrong 2.64%, 5.3 missing 3.75/wrong 8.59): **missing 6.67%,
wrong 9.01%, extra 0.96%** — the only family member regressing on
both axes; U+0670 convention effect 0.04pp. Bootstrap vs GLM-5.2:
+7.945pp, CI [+7.516, +8.370], p<1e-4. Regression axis complete:
5.2 2.5060 -> 5.3-Flash 8.5721 -> 5.3 9.9760 -> 4.7-flash 13.0035
raw. Recorded: results dir + RESULTS.md section/ledger/bootstrap
rows; paper.adoc row; site note amended (interscript.github.io
#149).

## Protocol

Expand Down
10 changes: 6 additions & 4 deletions TODO.publish/04-site-frontier-sentence.md
Original file line number Diff line number Diff line change
@@ -1,9 +1,11 @@
# 04 — Site frontier sentence (interscript.org ml page)

Status: BLOCKED on TODO.publish/01 (needs the full regression axis).
The ml.astro page exists on interscript.org (rebrand 2026-08); what
it lacks is the one-line frontier-context claim that the GLM sweep
now substantiates.
Status: COMPLETE (2026-09-02). Frontier-context note merged on the
site ledger section (interscript.github.io #147), amended with the
fourth row glm-4.7-flash 13.00 raw DER (#149) — wording "the rest of
the family" since 4.7 is a predecessor generation. Links
rababa/docs/RESULTS.md for the disclosed rows. Site guard (34/34),
astro check, full unit suite (255/255) green on both PRs.

## The claim (draft, number slots to fill from 01)

Expand Down
11 changes: 6 additions & 5 deletions TODO.training-work/02-glm-4-7-flash-row-completion.md
Original file line number Diff line number Diff line change
@@ -1,10 +1,11 @@
# 02 — GLM-4.7-Flash row completion

Status: IN FLIGHT (2026-09-01). Old API semantics verified by probe:
thinking.type=disabled ACCEPTED (plain completion expressible), but
the endpoint sits behind sustained 429s (code 1305) — running at
GLM_WORKERS=1, ~2 rows/min aggregate; ETA overnight. Checkpoint
resumes; empties self-heal (#65).
Status: COMPLETE (2026-09-02) — 1,200/1,200 after 12 resumable
passes past sustained 429s (sentinels 249 -> ... -> 0). 13.0035 raw
/ 13.2256 zero-skip; attribution missing 6.67 / wrong 9.01 / extra
0.96; bootstrap vs GLM-5.2 +7.945pp CI [+7.516, +8.370]. RESULTS +
ledger + paper row + site amendment landed. Full record:
TODO.publish/01.

## Remaining steps

Expand Down
17 changes: 17 additions & 0 deletions docs/RESULTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -433,6 +433,21 @@ behind 5.2 (CI [-18,028, -14,685] error positions). The dedicated-
model thesis gains its cleanest form: the newest generalist
generation lost classical-Arabic mark knowledge its predecessor had.

## glm-4.7-flash on SadeedDiac-25 (2026-09-02)

**13.2256/10.3206 zero-skip, 13.0035/10.0510 raw** (Misraj;
`results/sadeed-glm-4-7-flash/`; 1,200/1,200, zero empties — the
endpoint sat behind sustained 429s and needed 12 resumable passes,
sentinels dropping 249 -> ... -> 0). thinking-disabled accepted:
the last plain-completion GLM row. Worst frontier DER, and the only
family member regressing on both axes (missing 6.67% — worst of the
family, NFDW 17.6 — plus wrong 9.01%); U+0670 convention effect
0.04pp. Bootstrap vs GLM-5.2: +7.945pp, CI [+7.516, +8.370],
p<1e-4. The regression axis is complete: GLM-5.2 2.5060 raw ->
GLM-5.3-Flash 8.5721 -> GLM-5.3 9.9760 -> glm-4.7-flash 13.0035 —
every successor generation lost classical-Arabic mark knowledge its
predecessor had.

## Hebrew s46 on the Dicta test corpora (2026-09-01)

The disclosed modern-text gap, measured (eval_hebrew_dicta.py, TODO
Expand Down Expand Up @@ -496,6 +511,7 @@ row here does not belong in the paper.
| GLM-5.2 (2.5060 raw / 2.6911 zero-skip) | same evaluator | temp 0, `thinking.type=disabled` (plain completion) | raw: word-structure skips; zero-skip: projected | yes — script + CSVs (results/sadeed-glm-5-2/) |
| GLM-5.3-Flash (8.5721 raw / 8.7978 zero-skip) | same evaluator | temp 0, `reasoning_effort=low` (thinking CANNOT be disabled — API 400 code 1210) | same; dagger-alif convention skips; 0 empty responses after sentinel purge | yes — script + CSVs (results/sadeed-glm-5-3-flash/) |
| GLM-5.3 (9.9760 raw / 9.8971 zero-skip) | same evaluator | temp 0, `reasoning_effort=low` (same 400/1210 rejection) | same; 41 sentinels caught + refetched, 0 empties final | yes — script + CSVs (results/sadeed-glm-5-3/) |
| glm-4.7-flash (13.0035 raw / 13.2256 zero-skip) | same evaluator | temp 0, `thinking.type=disabled` (still accepted on 4.x) | same; 12 passes past sustained 429s, sentinels 249 -> ... -> 0, 0 empties final | yes — script + CSVs (results/sadeed-glm-4-7-flash/) |
| Claude-3.7-Sonnet (1.3941) | published number | their protocol, undisclosed to us | unknown | no — vendor-published |
| Gemini-Flash-2.0 (3.1926) / GPT-4 (3.8645) | published numbers | theirs | unknown | no — vendor-published |
| Sadeed-1.5B (7.2915) | published, same benchmark | theirs (their repo reports 1.2 under its own split — not comparable) | unknown | partially — paper + code, split differs |
Expand All @@ -519,6 +535,7 @@ identical.
| r7 vs r6 | −0.367 | [−0.467, −0.278] | <1e-4 — improvement is real |
| r7 vs GLM-5.2 (projected) | +0.539 (GLM worse) | [+0.202, +0.939] | 0.0003 — our teacher leads |
| GLM-5.3-Flash vs GLM-5.2 (projected) | +7.966 | [+7.520, +8.368] | <1e-4 — the frontier regression is overwhelming |
| glm-4.7-flash vs GLM-5.2 (projected) | +7.945 | [+7.516, +8.370] | <1e-4 — the axis endpoint; 4.7 regresses on both missing and wrong |

Policy (B1 of the improvement plan): every delta quoted in docs or
the paper carries its CI or an explicit note that the run predates
Expand Down
50 changes: 50 additions & 0 deletions results/sadeed-glm-4-7-flash/README.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,50 @@
# glm-4.7-flash on SadeedDiac-25 (2026-09-02)

Same harness as every LLM row: neutral prompt, `temperature 0`,
structure-preserving cleanup, Misraj evaluator, 1,200/1,200 responses,
zero empty responses. Decode protocol: `thinking: {"type": "disabled"}`
— glm-4.7-flash still accepts disabled thinking (unlike the 5.x
generation, which rejects it with HTTP 400 code 1210), so this row is
the last plain-completion GLM row.

Provider note: the endpoint sat behind sustained 429s (code 1305) —
the full set took 12 resumable passes (the #65 guard dropped
exhausted-retry empty sentinels between passes: 249 -> 208 -> 79 ->
37 -> 20 -> 13 -> 6 -> 4 -> 2 -> 0). All 1,200 final rows are real
responses.

## Results (Misraj evaluator, percentages)

| Protocol | DER (CE) | DER (w/o CE) | WER (CE) | WER (w/o CE) | NFDW |
|---|---|---|---|---|---|
| raw | **13.0035** | 10.0510 | 32.4140 | 26.7750 | 17.60 |
| projected zero-skip | **13.2256** | 10.3206 | 36.0985 | 30.1002 | 18.56 |

## Attribution (same 1,200 paragraphs, convention-normalized, rates over GT-marked positions)

| model | missing | **wrong haraqat** | extra | protocol |
|---|---|---|---|---|
| our r7 teacher | 0.11% | **2.62%** | 0.15% | greedy argmax |
| GLM-5.2 | 0.61% | **2.64%** | 0.17% | thinking disabled |
| GLM-5.3-Flash | 1.01% | **10.05%** | 0.20% | reasoning_effort=low |
| GLM-5.3 (full) | 3.75% | **8.59%** | 0.11% | reasoning_effort=low |
| glm-4.7-flash | **6.67%** | **9.01%** | 0.96% | thinking disabled |

glm-4.7-flash is the worst frontier row on DER and the only one that
regresses on BOTH axes simultaneously: the highest missing rate of
the family (6.67% — it under-diacritizes, NFDW 17.6%) alongside a
5.3-class wrong rate (9.01%). The U+0670 convention effect is
0.04pp — negligible, consistent with the family. Paired bootstrap vs
GLM-5.2 (paragraph-level simplified scorer, 10,000 resamples, seed
42): **+7.945pp, 95% CI [+7.516, +8.370], one-sided p<1e-4** —
decisively behind 5.2, whose plain-completion DER was 2.5060 raw /
2.6911 zero-skip.

The regression axis is now complete: GLM-5.2 2.51 -> GLM-5.3-Flash
8.57 -> GLM-5.3 9.98 -> glm-4.7-flash 13.00 raw DER. Every successor
generation lost classical-Arabic mark knowledge its predecessor had;
the dedicated-model thesis holds across the whole measurable family.

## Files

- `sadeed_preds_raw.csv`, `sadeed_preds_projected.csv`
Loading