Skip to content

FunctionGemma export: convert MLIR-text rewrites to StableHLO/DAG transforms #248

Description

@michalharakal

The FunctionGemma KV/int8 decode graphs are DSL/DAG-authored, but the graph post-processing in
FunctionGemmaExport.kt is done by regex rewrites of the emitted StableHLO text, not graph passes.
Convert each to the SKaiNET way (a StableHLO/DAG transform):

  • bf16 weight emission (rewriteGlobalsToBf16) → a DAG/StableHLO dtype transform (leverage DtypeForwardPropagationPass).
  • int8 quant (rewriteGlobalsToInt8 injecting dequant as text) → a real graph quantization pass; keep the host per-row quant writer, emit i8→f32 × scale as graph nodes.
  • Dynamic KV dims — retire the sentinel-prime 7919 → x?x string relax (most fragile: a magic prime that must not collide with a real dim/SSA id) → proper dynamic-shape tracing.
  • GemmaModel.forwardWithPast attnWithPast (hand-wired single-token attention) → an MHA-with-past module forward.
  • refsFor positional sub-module resolution → typed HybridTransformerBlock field access.

Context: PR #245. Tracker: SKaiNET-embedded sl2610-function-calling/docs/GEMMA-KV-INT8.md.

Activity

  1. added a commit that references this issue on Aug 11, 2026
  2. michalharakal commented on Aug 11, 2026

    @michalharakal
    ContributorAuthor

    Noting for scope-tracking: PR #302 moves the FunctionGemma export (including the bf16/int8 MLIR-text rewrites this issue is about — rewriteGlobalsToBf16, rewriteGlobalsToInt8, the sentinel-prime GEMMA_SENTINEL_PAST rollback) from :llm-runtime:kgemma into the new :llm-inference:functiongemma module, verbatim (golden-equivalence gate: byte-identical MLIR/safetensors vs. the pre-move export). The rewrites themselves are untouched — converting them to real StableHLO/DAG transforms stays this issue's scope, just relocated to FunctionGemmaExportHarness in the new module.

    Also surfaced while docker-vmfb-parity-testing the moved export (unrelated to the text-rewrite conversion, but possibly of interest here): host-CPU (x64) iree-run-module on the compiled gemma/gemma_with_past graphs diverges from the board-verified oracle at the first generated special-vocabulary token, reproducing identically in bf16 and FP32 and independent of the KV-cache loop — looks like a stock-IREE llvm-cpu / embedding-gather behavior for high (added-vocabulary) token indices that the board's aarch64 NEON Torq-fork path doesn't hit. Pre-existing, not a regression from the move. Diagnostic trail is in FunctionGemmaVmfbParityTest's class doc in the new module.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions