refactor(apertus, gemma3n): SafeTensors loading collapses onto the engine loader (SKaiNET#1246) - #401
Merged
Conversation
…er (SKaiNET#1246) The sharded ApertusSafeTensorsLoader no longer hand-rolls per-tensor materialization: it is one ShardedSafeTensorsParametersLoader.withPolicy( indexPath, dtypePolicy, tensorFilter = family allowlist) run into a map by HF name, followed by the table-driven HF -> GGUF renaming. The family keeps only what is genuinely its own — the allowlist, the [1, dim] norm shape normalization (re-wrapped over the same buffer, no copy), the xIELU scalar extraction, and tied embeddings. loadAndConvertTensor, loadScalarParam, transposeRowMajor (dead — every call site passed transpose = false) and the DequantOps usage on this lane are gone; new dtypePolicy constructor parameter (default Any, source-compatible). ApertusSingleSafeTensorsLoader (single-file lane) is unchanged: the engine's single-file SafeTensorsParametersLoader has no tensorFilter yet, so it cannot reproduce the warn-and-skip semantics for non-float tensors. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…er (SKaiNET#1246) Gemma3nSafeTensorsWeightLoader no longer hand-rolls per-tensor materialization: one ShardedSafeTensorsParametersLoader.withPolicy( indexPath, dtypePolicy, tensorFilter = family allowlist + PLE size guard) run into a map by HF name, then table-driven HF -> GGUF renaming (required embed/norm, tied output, every other slot optional as before). loadAndConvertTensor, transposeRowMajor (dead — every call site passed transpose = false) and the DequantOps import are gone; new dtypePolicy constructor parameter (default Any, source-compatible). Behavior change worth knowing: a per_layer_token_embd table over the 2 GiB single-array ceiling is now skipped (PLE disabled downstream, as on the Gemma 4 lane) where the old code would have failed in loadTensorData. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
…afetensors-loaders
This was referenced Sep 2, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
#1246 Phase 2, step 5 (part 2 of 2) — the Apertus and Gemma 3n sharded loaders, following #398's recipe.
loadToMap=ShardedSafeTensorsParametersLoader.withPolicy(indexPath, dtypePolicy, tensorFilter = name ∈ wantedHfNames)→ table-driven renaming (10 required pairs/layer) + tied/lm_headhandling. Family-side only:[1, dim]norm normalization, xIELU scalar extraction,Missing required tensorerrors. DeletedloadAndConvertTensor,loadScalarParam,transposeRowMajor(dead), the byte-path norm helper; per-tensorprintlns gone.tensorFilter= allowlist + PLE size guard onper_layer_token_embd.weight(mirrors Gemma'sMAX_BYTES_PER_TENSOR); required embed/norm + tied output, everything else optional (25 pairs/layer, 5 optional globals). An oversized PLE table is now skipped (PLE disabled downstream) where the old code threw insideloadTensorData. Deleted the hand-rolled dequant and theDequantOpsimport.dtypePolicy: DTypePolicy = Any(call sites source-compatible). Neither had a live transpose.QUANT4+.qbformat and unmapped tensors to skip — needs the single-filetensorFilter(SafeTensors: tensorFilter on the single-file SafeTensorsParametersLoader (parity with the sharded loader) SKaiNET#1256).Testing (published 0.53.0):
ApertusSafeTensorsLoaderFixtureTest(2) andGemma3nSafeTensorsWeightLoaderFixtureTest(2) — slots, norm normalization, xIELU floats, tied embedding, INT64 decoy exempt from the pre-scan,Require(BF16)→Bf16DenseTensorData;:llm-inference:apertus:jvmTest20/0,:llm-inference:gemma3n:jvmTest16/0; kapertus + kgemma3n compile.