Skip to content

feat(gemma3n)!: the DSL lane — AltUp/Laurel/sparsity/PLE, parity-gated vs llama.cpp (#377) - #393

Merged
michalharakal merged 1 commit into
developfrom
feat/gemma3n-dsl
Sep 2, 2026
Merged

michalharakal merged 1 commit into
developfrom
feat/gemma3n-dsl

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

The #377 DSL migration, per the maintainer decision: gemma3n gets a proper DSL definition — running eagerly like every family, and traceable for the StableHLO → IREE mobile path (next PR).

The hand-rolled runtime was never faithful

Recon against the real gemma-3n-E2B-it GGUF + HF modeling_gemma3n.py (read from source):

  • Gemma3nRuntime loaded the PLE tensors but never applied them (runLayer/forward never read them), had no Laurel block, and its AltUp ignored the modality router entirely (host-side coefficient math).
  • E2B_DEFAULT claimed AltUp/sparsity are E4B-only — the real E2B has altup.num_inputs=4 and first-10-layer activation sparsity (activation_sparsity_scale = Φ⁻¹(0.95)=1.6449 ×10, then −inf ×20).
  • Gate was 0/5 for a reason.

The DSL lane (faithful to HF, all through ctx.ops)

Piece What it does
Gemma3nAltUpBlock predict/correct with the tanh modality router; per-layer coef linears; correct_output_scale
Gemma3nAltUpGlobals stream init/merge over the 3D altup_proj/altup_unembd_proj stacks with per-token magnitude renorm (ggml ne-order → row-major reinterpret handled)
Gemma3nLaurelBlock x + post_laurel_norm(right(left(x)))
Gemma3nSparseGeGluFFN in-graph Gaussian-top-k on gate activations (population std; per-layer multipliers from the GGUF; −inf = off)
Gemma3nPerLayerApply PLE delta added to the non-active streams; reuses gemma's PerLayerEmbedding (math verified identical to HF)
Gemma3nModel GemmaModel-pattern orchestrator: predict → attn+Laurel with 1/√2 gating → sparse FFN → correct → PLE, per layer; per-type shared KV (last 10) via OwnerReadOnlyKVCache; hybrid sliding/global attention, dual RoPE bases, qk-norm + parameterless v-norm, attention scale 1.0

Loading stays packed/MAPPED through the engine (the 262k×7680 PLE table row-dequants on demand). Metadata/parser now read the real llama.cpp GGUF keys (sliding_window_pattern booleans, per-layer activation_sparsity_scale, rope.freq_base fallback, rms_norm_eps, per-layer feed_forward_length). Strict binding both directions (the qwen-bias lesson).

Gate — the last ungated generative family

Gemma3nGoldenTokenParityTest: full 32-step greedy text equality vs mainline llama.cpp b10621 on gemma-3n-E2B-it-Q4_K_M.gguf — GREEN. ("The capital of France is" divergence was measured as a 0.12-nat top-2 near-tie at step 3, i.e. Q4_K cross-implementation noise, not a defect; the fixture prompt has a 1.77-nat worst margin.) Wired into smoke-reference (gemma3n_gguf_url + 20g heap arg); smoke-models.json gains a Gemma3n-E2B row.

Routing

kgemma (GEMMA3N/GGUF) and skainet-cli (gemma3n arch, refusal removed) both run the DSL lane; verified end-to-end — both decode " 2, 3, 5, 7, …". SafeTensors stays on the legacy runtime until the DSL grows that leg (tracked).

Verification

Parity gate green; gemma3n/gemma/kgemma/kgemma3n jvmTest + apiCheck green (dump regenerated); macosArm64/js/wasmWasi compile; both CLIs verified against the real model.

With this, every shipped generative family has a reference-implementation parity gate.

Next PR (in flight): exportGemma3n — StableHLO redecode graph + manifest for the IREE/Android mobile path, plus the gemma3n docs page.

🤖 Generated with Claude Code

…d vs llama.cpp (#377)

Gemma 3n gets a proper DSL definition and stops depending on the
hand-rolled Gemma3nRuntime for GGUF inference. The recon showed that
runtime was never faithful to real checkpoints: it loaded the PLE
tensors but never applied them, had no Laurel block, ignored the AltUp
modality router (host-side coefficient math), and its E2B_DEFAULT
claimed AltUp/sparsity are E4B-only — the real gemma-3n-E2B GGUF has
altup.num_inputs=4 and first-10-layer activation sparsity. Gate was
0/5 for a reason.

The new lane, faithful to HF modeling_gemma3n.py (read from source):

- Gemma3nAltUpBlock: predict/correct with the tanh modality router
  (router_norm * 1/hidden -> modality_router -> tanh), per-layer
  coefficient linears, correct_output_scale.
- Gemma3nAltUpGlobals: stream init/merge over the GGUF's 3D
  altup_proj/altup_unembd_proj stacks with per-token magnitude
  renormalization (the ggml ne-order vs row-major reinterpret is
  handled by reshaping to [numExtra, H, H] before slicing).
- Gemma3nLaurelBlock; Gemma3nSparseGeGluFFN (in-graph Gaussian-top-k,
  population std, per-layer std multipliers from the GGUF, -inf = off);
  Gemma3nPerLayerApply (PLE delta added to the NON-active streams,
  reusing gemma's PerLayerEmbedding whose math is identical to HF).
- gemma3nNetwork() + Gemma3nModel: GemmaModel-pattern orchestrator
  threading 4 AltUp streams per layer (predict -> attn+Laurel with
  1/sqrt2 gating -> sparse FFN -> correct -> PLE), per-type shared KV
  for the last 10 layers via OwnerReadOnlyKVCache, hybrid
  sliding/global attention with dual RoPE bases, qk-norm +
  parameterless v-norm, attention scale 1.0. Everything through
  ctx.ops — traceable for the StableHLO/IREE mobile path (next PR).
- Metadata/parser now read the real llama.cpp GGUF keys:
  sliding_window_pattern booleans, per-layer activation_sparsity_scale
  (1.6449/-inf encoding), rope.freq_base fallback (1M global, 10k SWA
  default), rms_norm_eps, per-layer feed_forward_length.
- Gemma3nNetworkLoader: engine loading stays packed/MAPPED (PLE table
  row-dequants on demand), strict binding both directions.
- Routing: kgemma GEMMA3N/GGUF and skainet-cli gemma3n now run the DSL
  lane; SafeTensors stays on the legacy runtime until the DSL grows
  that leg.

Gate (the last ungated generative family): Gemma3nGoldenTokenParityTest
asserts full 32-step greedy text equality vs mainline llama.cpp b10621
on gemma-3n-E2B-it-Q4_K_M — GREEN. Fixture prompt chosen for greedy
decisiveness (min top-1/top-2 margin 1.77 nats; "The capital of France
is" hits a measured 0.12-nat tie at step 3 that Q4_K noise flips).
Wired into smoke-reference (gemma3n_gguf_url + 20g heap arg);
smoke-models.json gains a Gemma3n-E2B row. Both CLIs decode
" 2, 3, 5, 7, ..." end-to-end.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit 18c6b2d into develop Sep 2, 2026
2 checks passed
@michalharakal
michalharakal deleted the feat/gemma3n-dsl branch September 2, 2026 08:42
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant