From 67047dbb109ca0aa01cc87e326f9b734550594e2 Mon Sep 17 00:00:00 2001 From: Ronald Tse Date: Tue, 1 Sep 2026 06:30:02 +0200 Subject: [PATCH] =?UTF-8?q?docs:=20GLM-5.3-Flash=20attribution=20=E2=80=94?= =?UTF-8?q?=20the=20regression=20is=20wrong=20haraqat,=20not=20convention?= MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Corrects the earlier driver claim. Per-position decomposition over the same 1,200 paragraphs, convention-normalized with U+0670 rules derived from the aligned positions (drop after ى, fatha elsewhere; controls r7/G-5.2 move <=0.009pp): GLM-5.3-Flash wrong-haraqat rate 10.05% vs GLM-5.2's 2.64% — which matches our r7 teacher's 2.62% to within 0.02pp — with missing 1.01%, extra 0.20%, and the entire dagger-alif convention effect at 0.125pp (310 marks, zero in GT). The regression is genuine mark errors at ~4x the predecessor rate; the Quranic marks explain the raw-mode skip flood, not the DER gap. --- docs/RESULTS.md | 19 ++++++++++----- results/sadeed-glm-5-3-flash/README.md | 32 +++++++++++++++++--------- 2 files changed, 34 insertions(+), 17 deletions(-) diff --git a/docs/RESULTS.md b/docs/RESULTS.md index 9387634..5286fde 100644 --- a/docs/RESULTS.md +++ b/docs/RESULTS.md @@ -66,12 +66,19 @@ protocol, zero skipped paragraphs. REJECTS disabled thinking (HTTP 400 code 1210) — reasoning cannot be turned off, only dialed to low/high/max — so this ran at `reasoning_effort=low`, the nearest analog to plain completion. - ~3.4x worse DER than GLM-5.2 and behind Sadeed-1.5B, driven by - Quranic-convention orthography (dagger-alif forms the GT does not - use, 9.8% not-fully-diacritized words) — a generalist regression on - classical-knowledge output conventions, recorded as a dedicated-model - data point. Zero empty responses after purging 140 checkpoint rows - the pre-fix payload had retried into empty strings. + ~3.4x worse DER than GLM-5.2 and behind Sadeed-1.5B. Attribution + (2026-09-01, per-position decomposition over the same 1,200 + paragraphs, convention-normalized): **wrong-haraqat rate 10.05% vs + GLM-5.2's 2.64%** — which matches our r7 teacher's 2.62% almost + exactly — while missing is 1.01%, extra 0.20%, and the entire + dagger-alif (U+0670) convention effect is 0.125pp (310 marks, zero + in GT; rules derived from the aligned positions: drop after ى, + fatha elsewhere; controls r7/G-5.2 move <=0.009pp). The regression + is genuine mark errors, not orthographic convention — the frontier + generalist lost core classical-Arabic knowledge its predecessor + had. The dagger-alif observation explains the raw-mode skip flood, + not the DER gap. Zero empty responses after purging 140 checkpoint + rows the pre-fix payload had retried into empty strings. - **r3 = r2 + 1 epoch on the decontaminated Misraj corpus (1M lines) + 150k MSA replay**: best non-frontier-LLM result on the benchmark; diff --git a/results/sadeed-glm-5-3-flash/README.md b/results/sadeed-glm-5-3-flash/README.md index 783de1e..ad83ebc 100644 --- a/results/sadeed-glm-5-3-flash/README.md +++ b/results/sadeed-glm-5-3-flash/README.md @@ -23,17 +23,27 @@ For reference, GLM-5.2 (thinking disabled): raw 2.5060/1.5537/7.9929, zero-skip 2.6911/1.7179/8.3037. GLM-5.3-Flash is ~3.4x worse on DER than its predecessor under the nearest equivalent protocol. -## Why the delta is real but protocol-shaped - -- **Orthography**: the model emits Quranic-convention marks (dagger - alif: عَلَىٰ, هٰذِهِ, ذَٰلِكَ; also بِهِۦ) where the benchmark's GT uses - plain MSA forms. The evaluator skips whole sentences on word - mismatches (survivorship in raw mode); surviving dagger-alif words - count as diacritic errors. Not-fully-diacritized words run 9.8%. -- **Thinking floor**: reasoning cannot be turned off, so the - plain-completion protocol the published LLM rows used is not - expressible for this model; `low` still engaged reasoning on long - paragraphs (reasoning_content observed in-flight). +## Why the delta is real (attribution, 2026-09-01) + +Per-position decomposition over the same 1,200 paragraphs, +convention-normalized (U+0670 rules derived from aligned positions: +drop after ى, fatha on other letters; controls move <=0.009pp): + +| model | missing | **wrong haraqat** | extra | U+0670 convention | +|---|---|---|---|---| +| our r7 teacher | 0.11% | **2.62%** | 0.15% | — | +| GLM-5.2 | 0.61% | **2.64%** | 0.17% | ~0.009pp | +| GLM-5.3-Flash | 1.01% | **10.05%** | 0.20% | 0.125pp | + +The regression is overwhelmingly WRONG haraqat at ~4x its +predecessor's rate (which matches our dedicated teacher's to within +0.02pp) — not the dagger-alif orthography (0.125pp total), not +under-diacritization (1.01%), not thinking overhead. The Quranic +marks (عَلَىٰ, هٰذِهِ, ذَٰلِكَ; 310 in the outputs, zero in GT) explain the +raw-protocol evaluator skips, not the DER gap. +- **Thinking floor** (protocol caveat, unchanged): reasoning cannot be + turned off; `low` still engaged reasoning on long paragraphs + (reasoning_content observed in-flight). - **Measurement hygiene**: an initial run resumed from a 140-row checkpoint whose rows were empty responses produced by the pre-fix both-knobs payload (HTTP 400s retried into empty strings) —