feat(gemma3n)!: the DSL lane — AltUp/Laurel/sparsity/PLE, parity-gated vs llama.cpp (#377) - #393
Merged
Merged
Conversation
…d vs llama.cpp (#377) Gemma 3n gets a proper DSL definition and stops depending on the hand-rolled Gemma3nRuntime for GGUF inference. The recon showed that runtime was never faithful to real checkpoints: it loaded the PLE tensors but never applied them, had no Laurel block, ignored the AltUp modality router (host-side coefficient math), and its E2B_DEFAULT claimed AltUp/sparsity are E4B-only — the real gemma-3n-E2B GGUF has altup.num_inputs=4 and first-10-layer activation sparsity. Gate was 0/5 for a reason. The new lane, faithful to HF modeling_gemma3n.py (read from source): - Gemma3nAltUpBlock: predict/correct with the tanh modality router (router_norm * 1/hidden -> modality_router -> tanh), per-layer coefficient linears, correct_output_scale. - Gemma3nAltUpGlobals: stream init/merge over the GGUF's 3D altup_proj/altup_unembd_proj stacks with per-token magnitude renormalization (the ggml ne-order vs row-major reinterpret is handled by reshaping to [numExtra, H, H] before slicing). - Gemma3nLaurelBlock; Gemma3nSparseGeGluFFN (in-graph Gaussian-top-k, population std, per-layer std multipliers from the GGUF, -inf = off); Gemma3nPerLayerApply (PLE delta added to the NON-active streams, reusing gemma's PerLayerEmbedding whose math is identical to HF). - gemma3nNetwork() + Gemma3nModel: GemmaModel-pattern orchestrator threading 4 AltUp streams per layer (predict -> attn+Laurel with 1/sqrt2 gating -> sparse FFN -> correct -> PLE), per-type shared KV for the last 10 layers via OwnerReadOnlyKVCache, hybrid sliding/global attention with dual RoPE bases, qk-norm + parameterless v-norm, attention scale 1.0. Everything through ctx.ops — traceable for the StableHLO/IREE mobile path (next PR). - Metadata/parser now read the real llama.cpp GGUF keys: sliding_window_pattern booleans, per-layer activation_sparsity_scale (1.6449/-inf encoding), rope.freq_base fallback (1M global, 10k SWA default), rms_norm_eps, per-layer feed_forward_length. - Gemma3nNetworkLoader: engine loading stays packed/MAPPED (PLE table row-dequants on demand), strict binding both directions. - Routing: kgemma GEMMA3N/GGUF and skainet-cli gemma3n now run the DSL lane; SafeTensors stays on the legacy runtime until the DSL grows that leg. Gate (the last ungated generative family): Gemma3nGoldenTokenParityTest asserts full 32-step greedy text equality vs mainline llama.cpp b10621 on gemma-3n-E2B-it-Q4_K_M — GREEN. Fixture prompt chosen for greedy decisiveness (min top-1/top-2 margin 1.77 nats; "The capital of France is" hits a measured 0.12-nat tie at step 3 that Q4_K noise flips). Wired into smoke-reference (gemma3n_gguf_url + 20g heap arg); smoke-models.json gains a Gemma3n-E2B row. Both CLIs decode " 2, 3, 5, 7, ..." end-to-end. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The #377 DSL migration, per the maintainer decision: gemma3n gets a proper DSL definition — running eagerly like every family, and traceable for the StableHLO → IREE mobile path (next PR).
The hand-rolled runtime was never faithful
Recon against the real
gemma-3n-E2B-itGGUF + HFmodeling_gemma3n.py(read from source):Gemma3nRuntimeloaded the PLE tensors but never applied them (runLayer/forwardnever read them), had no Laurel block, and its AltUp ignored the modality router entirely (host-side coefficient math).E2B_DEFAULTclaimed AltUp/sparsity are E4B-only — the real E2B hasaltup.num_inputs=4and first-10-layer activation sparsity (activation_sparsity_scale= Φ⁻¹(0.95)=1.6449 ×10, then −inf ×20).The DSL lane (faithful to HF, all through
ctx.ops)Gemma3nAltUpBlockcorrect_output_scaleGemma3nAltUpGlobalsaltup_proj/altup_unembd_projstacks with per-token magnitude renorm (ggml ne-order → row-major reinterpret handled)Gemma3nLaurelBlockx + post_laurel_norm(right(left(x)))Gemma3nSparseGeGluFFNGemma3nPerLayerApplyPerLayerEmbedding(math verified identical to HF)Gemma3nModelOwnerReadOnlyKVCache; hybrid sliding/global attention, dual RoPE bases, qk-norm + parameterless v-norm, attention scale 1.0Loading stays packed/MAPPED through the engine (the 262k×7680 PLE table row-dequants on demand). Metadata/parser now read the real llama.cpp GGUF keys (
sliding_window_patternbooleans, per-layeractivation_sparsity_scale,rope.freq_basefallback,rms_norm_eps, per-layerfeed_forward_length). Strict binding both directions (the qwen-bias lesson).Gate — the last ungated generative family
Gemma3nGoldenTokenParityTest: full 32-step greedy text equality vs mainline llama.cpp b10621 ongemma-3n-E2B-it-Q4_K_M.gguf— GREEN. ("The capital of France is" divergence was measured as a 0.12-nat top-2 near-tie at step 3, i.e. Q4_K cross-implementation noise, not a defect; the fixture prompt has a 1.77-nat worst margin.) Wired into smoke-reference (gemma3n_gguf_url+ 20g heap arg);smoke-models.jsongains a Gemma3n-E2B row.Routing
kgemma (GEMMA3N/GGUF) and skainet-cli (gemma3n arch, refusal removed) both run the DSL lane; verified end-to-end — both decode
" 2, 3, 5, 7, …". SafeTensors stays on the legacy runtime until the DSL grows that leg (tracked).Verification
Parity gate green; gemma3n/gemma/kgemma/kgemma3n jvmTest + apiCheck green (dump regenerated); macosArm64/js/wasmWasi compile; both CLIs verified against the real model.
With this, every shipped generative family has a reference-implementation parity gate.
Next PR (in flight):
exportGemma3n— StableHLO redecode graph + manifest for the IREE/Android mobile path, plus the gemma3n docs page.🤖 Generated with Claude Code