diff --git a/CHANGELOG.md b/CHANGELOG.md index 4d21458d..68b3c3ff 100644 --- a/CHANGELOG.md +++ b/CHANGELOG.md @@ -9,8 +9,40 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0 ## [Unreleased] +## [0.57.0] — 2026-09-25 + +Lock-step with **SKaiNET engine 0.57.0**. Headline: **Qwen on the compiled IREE KV path**: the +qwen-kv-v1 export, runtime and templates, with Qwen3-0.6B verified token-for-token against llama.cpp +on host IREE; on-device validation is still to come (see the entry below for exact status). + +### Changed + +- **Engine 0.57.0, Kotlin 2.4.20** (#453). The engine fixes the StableHLO export of an explicit + attention mask under grouped-query attention and onto a dynamic key length (SKaiNET#1302), and + makes kotlinx-io part of `skainet-data-source`'s API. This repo moves to Kotlin 2.4.20 with it. +- **Dependencies**: Ktor client 3.6.0 (#448), kctfork 0.14.0 (#449), jackson-databind 2.22.3 (#454). + +### Added + +- **FunctionGemma export contract carries the embedding geometry** (#452). `FunctionGemmaSpec` gains + `vocabSize` (default 262144), `nHeads`, `hiddenSize` and `slidingWindow`, and `manifest.json` emits + them, so `IreeKvSpec.fromManifest` no longer falls back to the stock constants. The export CLI reads + the vocabulary size from the checkpoint's `token_embd.weight` (`GEMMA_VOCAB` overrides): a fine-tune + that added special tokens has more rows, and the native KV session locates the embedding table by + `vocabSize × hiddenSize` bytes and clamps ids ≥ `vocabSize` to 0, so with the stock constant those + tokens would silently have become token 0. `IreeKvSpec.functionGemma270m` takes `vocabSize`. + +### Fixed + +- **Stale `functiongemma` API dump.** #452 changed the public API without re-dumping it, so + `apiCheck` failed on `develop`; refreshed in #453. + ### Added — Qwen on the compiled IREE KV path (closes #411, #409) +Status in this release: Qwen3-0.6B exports, compiles for Vulkan (valhall4) and arm32, and matches +mainline llama.cpp's 32 greedy tokens on host IREE (`QwenVmfbParityTest`). Not yet run on a device. +Qwen2.5-0.5B export is blocked by a SKaiNET core gap (attention bias does not externalize). + - **`IreeKvSession` / `IreeKvSpec` generalized for GQA and single-RoPE-base models** (`llm-runtime:iree-android`): `nKvHeads` may now be any value with `nHeads % nKvHeads == 0` (grouped-query attention — Qwen2.5-0.5B nKvHeads=2, Qwen3-0.6B nKvHeads=8), not just @@ -98,15 +130,6 @@ A transformers-only release against **SKaiNET engine 0.56.0** (unchanged). ### Added -- **`IreeMoonshineStream`** (`llm-runtime:iree-android`, `libskainet_moonshine_stream.so`, arm64-v8a + - armeabi-v7a, Vulkan + local-task): the streaming Moonshine v2 speech-to-text runtime over the five - graphs of `MoonshineV2ExportCli` — PCM in, cumulative partial transcripts out, exact final on - `finish()`. The counterpart of `IreeKvSession` for ASR: the last piece a Moonshine cartridge needed - from a released artifact. `native/build-moonshine-stream.sh` builds it with the same image and - cache as the other two libraries. - -### Added - - **`IreeMoonshineStream`** (`llm-runtime:iree-android`, `libskainet_moonshine_stream.so`, arm64-v8a + armeabi-v7a, Vulkan + local-task): the streaming Moonshine v2 speech-to-text runtime over the five graphs of `MoonshineV2ExportCli` — PCM in, cumulative partial transcripts out, exact final on diff --git a/README.md b/README.md index a58a01b0..9aa857c4 100644 --- a/README.md +++ b/README.md @@ -109,18 +109,24 @@ Honest status — see the project-status note at the top of this README. ## Current release -The current release is **0.56.2** (against **SKaiNET 0.56.0** — a transformers-only release, same -pattern as 0.54.1: no new engine version needed). - -**`IreeMoonshineStream`: the streaming Moonshine v2 speech-to-text runtime.** `libskainet_moonshine_stream.so` -(`llm-runtime:iree-android`, arm64-v8a + armeabi-v7a, Vulkan or CPU) drives the five graphs of -`MoonshineV2ExportCli` — frontend, encoder, adapter, masked prefill, dynamic with-past step — plus the -shared decoder parameter archive as a real streaming loop on the device: PCM in, cumulative partial -transcripts out, one exact full re-decode on `finish()`. The Kotlin binding is shaped like `IreeKvSession`: -files by absolute path, no `Context`. With the exporter (0.56.1) this is the last piece a Moonshine -cartridge needed from a released artifact rather than from C source of its own. - -It builds on **0.56.1**, which released `MoonshineV2ExportCli` (a Hugging Face snapshot in, the five +The current release is **0.57.0**, in lock-step with **SKaiNET 0.57.0**. + +**Qwen on the compiled IREE KV path.** `QwenExportHarness` / `QwenExportCli` trace Qwen2 and Qwen3 +checkpoints to the three graphs of the `qwen-kv-v1` contract (catalog-prefix prefill, one-call chunk, +dynamic-cache decode step), `IreeKvSession` runs grouped-query models, and `Qwen25ChatTemplate` renders +Qwen2.5's official tool-calling template. Qwen3-0.6B compiles for Vulkan and arm32 and reproduces +llama.cpp's greedy tokens exactly on host IREE; it has not been run on a device yet, and Qwen2.5 export +waits on an engine fix for attention bias. **The Moonshine v2 stream** caps its final re-decode at the +model card's length budget, so short commands finish sooner. + +**Engine 0.57.0 and Kotlin 2.4.20.** The engine fixes the StableHLO export of an explicit attention mask +under grouped-query attention and onto a dynamic key length, the shape of a chunked-prefill graph over a +KV cache. **The FunctionGemma export contract now carries the embedding geometry** (`vocabSize`, +`nHeads`, `hiddenSize`, `slidingWindow`) into `manifest.json`, read from the checkpoint, so a fine-tune +with added special tokens exports and runs without its new tokens silently collapsing to token 0. + +It builds on **0.56.2**, which released `IreeMoonshineStream`, the streaming Moonshine v2 runtime on the +device; **0.56.1**, which released `MoonshineV2ExportCli` (a Hugging Face snapshot in, the five StableHLO graphs and two host-side tables out, no Python, nothing to configure per language or size), and **0.56.0**, which put the two version lines back in lock-step with **SKaiNET 0.56.0** (the engine skipped 0.55.0 to realign them). @@ -256,7 +262,7 @@ The recommended way to consume is via the BOM. It pins every published `skainet- ```kotlin dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) // Versions resolved from the BOM: implementation("sk.ainet.transformers:skainet-transformers-core") diff --git a/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc b/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc index 13d6a4f6..28ce7687 100644 --- a/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc +++ b/docs/modules/ROOT/pages/how-to/export-moonshine-v2-streaming.adoc @@ -32,7 +32,7 @@ From your own build, resolve the published module and run its main class: ---- val exportTool by configurations.creating dependencies { - exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) exportTool("sk.ainet.transformers:skainet-transformers-inference-moonshine") } tasks.register("exportMoonshine") { diff --git a/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc b/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc index ba4796f3..b22ea1c8 100644 --- a/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc +++ b/docs/modules/ROOT/pages/reference/moonshine-encoder.adoc @@ -24,7 +24,7 @@ a DSL decoder is future work. [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) implementation("sk.ainet.transformers:skainet-transformers-inference-moonshine") } ---- diff --git a/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc b/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc index 81d6f00e..562442fd 100644 --- a/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc +++ b/docs/modules/ROOT/pages/tutorials/android-getting-started.adoc @@ -55,7 +55,7 @@ dependencies { // self-registers it on ART at process start — nothing to call. runtimeOnly("sk.ainet.core:skainet-backend-jni-cpu") - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) implementation("sk.ainet.transformers:skainet-transformers-core") implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama") implementation("sk.ainet.transformers:skainet-transformers-inference-llama") diff --git a/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc b/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc index 4cc34d51..dea25147 100644 --- a/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc +++ b/docs/modules/ROOT/pages/tutorials/getting-started-java.adoc @@ -25,7 +25,7 @@ In your `build.gradle.kts`: [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama") implementation("sk.ainet.transformers:skainet-transformers-agent") @@ -41,7 +41,7 @@ Or in Maven (Maven needs the `-jvm` classifier suffix on platform artifacts): sk.ainet.transformers skainet-transformers-bom - 0.56.2 + 0.57.0 pom import diff --git a/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc b/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc index f5414521..0e28d2c7 100644 --- a/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc +++ b/docs/modules/ROOT/pages/tutorials/getting-started-leaf.adoc @@ -34,7 +34,7 @@ encoder output to the advertised dimensionality. The runtime applies it automati [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) implementation("sk.ainet.transformers:skainet-transformers-providers") } ---- diff --git a/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc b/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc index 30b91b99..03cd064a 100644 --- a/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc +++ b/docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc @@ -52,7 +52,7 @@ The pieces you need live in three modules: [source,kotlin] ---- dependencies { - implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2")) + implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0")) implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama") implementation("sk.ainet.transformers:skainet-transformers-agent") diff --git a/gradle.properties b/gradle.properties index cc3cc253..c0e454f3 100644 --- a/gradle.properties +++ b/gradle.properties @@ -1,5 +1,5 @@ GROUP=sk.ainet.transformers -VERSION_NAME=0.56.2 +VERSION_NAME=0.57.0 POM_DESCRIPTION=SKaiNET-transformers