Skip to content

refactor(apertus, gemma3n): SafeTensors loading collapses onto the engine loader (SKaiNET#1246) - #401

Merged
michalharakal merged 3 commits into
developfrom
feat/1246-family-safetensors-loaders
Sep 2, 2026
Merged

michalharakal merged 3 commits into
developfrom
feat/1246-family-safetensors-loaders

Conversation

@michalharakal

Copy link
Copy Markdown
Contributor

#1246 Phase 2, step 5 (part 2 of 2) — the Apertus and Gemma 3n sharded loaders, following #398's recipe.

  • Apertus: loadToMap = ShardedSafeTensorsParametersLoader.withPolicy(indexPath, dtypePolicy, tensorFilter = name ∈ wantedHfNames) → table-driven renaming (10 required pairs/layer) + tied/lm_head handling. Family-side only: [1, dim] norm normalization, xIELU scalar extraction, Missing required tensor errors. Deleted loadAndConvertTensor, loadScalarParam, transposeRowMajor (dead), the byte-path norm helper; per-tensor printlns gone.
  • Gemma 3n: same recipe; tensorFilter = allowlist + PLE size guard on per_layer_token_embd.weight (mirrors Gemma's MAX_BYTES_PER_TENSOR); required embed/norm + tied output, everything else optional (25 pairs/layer, 5 optional globals). An oversized PLE table is now skipped (PLE disabled downstream) where the old code threw inside loadTensorData. Deleted the hand-rolled dequant and the DequantOps import.
  • Both gain dtypePolicy: DTypePolicy = Any (call sites source-compatible). Neither had a live transpose.
  • Voxtral not collapsed: single-file with a custom QUANT4 + .qb format and unmapped tensors to skip — needs the single-file tensorFilter (SafeTensors: tensorFilter on the single-file SafeTensorsParametersLoader (parity with the sharded loader) SKaiNET#1256).

Testing (published 0.53.0): ApertusSafeTensorsLoaderFixtureTest (2) and Gemma3nSafeTensorsWeightLoaderFixtureTest (2) — slots, norm normalization, xIELU floats, tied embedding, INT64 decoy exempt from the pre-scan, Require(BF16) → Bf16DenseTensorData; :llm-inference:apertus:jvmTest 20/0, :llm-inference:gemma3n:jvmTest 16/0; kapertus + kgemma3n compile.

michalharakal and others added 3 commits September 2, 2026 14:54
…er (SKaiNET#1246)

The sharded ApertusSafeTensorsLoader no longer hand-rolls per-tensor
materialization: it is one ShardedSafeTensorsParametersLoader.withPolicy(
indexPath, dtypePolicy, tensorFilter = family allowlist) run into a map by
HF name, followed by the table-driven HF -> GGUF renaming. The family keeps
only what is genuinely its own — the allowlist, the [1, dim] norm shape
normalization (re-wrapped over the same buffer, no copy), the xIELU scalar
extraction, and tied embeddings. loadAndConvertTensor, loadScalarParam,
transposeRowMajor (dead — every call site passed transpose = false) and
the DequantOps usage on this lane are gone; new dtypePolicy constructor
parameter (default Any, source-compatible).

ApertusSingleSafeTensorsLoader (single-file lane) is unchanged: the
engine's single-file SafeTensorsParametersLoader has no tensorFilter yet,
so it cannot reproduce the warn-and-skip semantics for non-float tensors.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…er (SKaiNET#1246)

Gemma3nSafeTensorsWeightLoader no longer hand-rolls per-tensor
materialization: one ShardedSafeTensorsParametersLoader.withPolicy(
indexPath, dtypePolicy, tensorFilter = family allowlist + PLE size guard)
run into a map by HF name, then table-driven HF -> GGUF renaming (required
embed/norm, tied output, every other slot optional as before).
loadAndConvertTensor, transposeRowMajor (dead — every call site passed
transpose = false) and the DequantOps import are gone; new dtypePolicy
constructor parameter (default Any, source-compatible).

Behavior change worth knowing: a per_layer_token_embd table over the
2 GiB single-array ceiling is now skipped (PLE disabled downstream, as on
the Gemma 4 lane) where the old code would have failed in loadTensorData.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
@michalharakal
michalharakal merged commit 567cadc into develop Sep 2, 2026
2 checks passed
@michalharakal
michalharakal deleted the feat/1246-family-safetensors-loaders branch September 2, 2026 13:26
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant