Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
41 changes: 32 additions & 9 deletions CHANGELOG.md
Original file line number Diff line number Diff line change
Expand Up @@ -9,8 +9,40 @@ and this project adheres to [Semantic Versioning](https://semver.org/spec/v2.0.0

## [Unreleased]

## [0.57.0] — 2026-09-25

Lock-step with **SKaiNET engine 0.57.0**. Headline: **Qwen on the compiled IREE KV path**: the
qwen-kv-v1 export, runtime and templates, with Qwen3-0.6B verified token-for-token against llama.cpp
on host IREE; on-device validation is still to come (see the entry below for exact status).

### Changed

- **Engine 0.57.0, Kotlin 2.4.20** (#453). The engine fixes the StableHLO export of an explicit
attention mask under grouped-query attention and onto a dynamic key length (SKaiNET#1302), and
makes kotlinx-io part of `skainet-data-source`'s API. This repo moves to Kotlin 2.4.20 with it.
- **Dependencies**: Ktor client 3.6.0 (#448), kctfork 0.14.0 (#449), jackson-databind 2.22.3 (#454).

### Added

- **FunctionGemma export contract carries the embedding geometry** (#452). `FunctionGemmaSpec` gains
`vocabSize` (default 262144), `nHeads`, `hiddenSize` and `slidingWindow`, and `manifest.json` emits
them, so `IreeKvSpec.fromManifest` no longer falls back to the stock constants. The export CLI reads
the vocabulary size from the checkpoint's `token_embd.weight` (`GEMMA_VOCAB` overrides): a fine-tune
that added special tokens has more rows, and the native KV session locates the embedding table by
`vocabSize × hiddenSize` bytes and clamps ids ≥ `vocabSize` to 0, so with the stock constant those
tokens would silently have become token 0. `IreeKvSpec.functionGemma270m` takes `vocabSize`.

### Fixed

- **Stale `functiongemma` API dump.** #452 changed the public API without re-dumping it, so
`apiCheck` failed on `develop`; refreshed in #453.

### Added — Qwen on the compiled IREE KV path (closes #411, #409)

Status in this release: Qwen3-0.6B exports, compiles for Vulkan (valhall4) and arm32, and matches
mainline llama.cpp's 32 greedy tokens on host IREE (`QwenVmfbParityTest`). Not yet run on a device.
Qwen2.5-0.5B export is blocked by a SKaiNET core gap (attention bias does not externalize).

- **`IreeKvSession` / `IreeKvSpec` generalized for GQA and single-RoPE-base models**
(`llm-runtime:iree-android`): `nKvHeads` may now be any value with `nHeads % nKvHeads == 0`
(grouped-query attention — Qwen2.5-0.5B nKvHeads=2, Qwen3-0.6B nKvHeads=8), not just
Expand Down Expand Up @@ -98,15 +130,6 @@ A transformers-only release against **SKaiNET engine 0.56.0** (unchanged).

### Added

- **`IreeMoonshineStream`** (`llm-runtime:iree-android`, `libskainet_moonshine_stream.so`, arm64-v8a +
armeabi-v7a, Vulkan + local-task): the streaming Moonshine v2 speech-to-text runtime over the five
graphs of `MoonshineV2ExportCli` — PCM in, cumulative partial transcripts out, exact final on
`finish()`. The counterpart of `IreeKvSession` for ASR: the last piece a Moonshine cartridge needed
from a released artifact. `native/build-moonshine-stream.sh` builds it with the same image and
cache as the other two libraries.

### Added

- **`IreeMoonshineStream`** (`llm-runtime:iree-android`, `libskainet_moonshine_stream.so`, arm64-v8a +
armeabi-v7a, Vulkan + local-task): the streaming Moonshine v2 speech-to-text runtime over the five
graphs of `MoonshineV2ExportCli` — PCM in, cumulative partial transcripts out, exact final on
Expand Down
32 changes: 19 additions & 13 deletions README.md
Original file line number Diff line number Diff line change
Expand Up @@ -109,18 +109,24 @@ Honest status — see the project-status note at the top of this README.

## Current release

The current release is **0.56.2** (against **SKaiNET 0.56.0** — a transformers-only release, same
pattern as 0.54.1: no new engine version needed).

**`IreeMoonshineStream`: the streaming Moonshine v2 speech-to-text runtime.** `libskainet_moonshine_stream.so`
(`llm-runtime:iree-android`, arm64-v8a + armeabi-v7a, Vulkan or CPU) drives the five graphs of
`MoonshineV2ExportCli` — frontend, encoder, adapter, masked prefill, dynamic with-past step — plus the
shared decoder parameter archive as a real streaming loop on the device: PCM in, cumulative partial
transcripts out, one exact full re-decode on `finish()`. The Kotlin binding is shaped like `IreeKvSession`:
files by absolute path, no `Context`. With the exporter (0.56.1) this is the last piece a Moonshine
cartridge needed from a released artifact rather than from C source of its own.

It builds on **0.56.1**, which released `MoonshineV2ExportCli` (a Hugging Face snapshot in, the five
The current release is **0.57.0**, in lock-step with **SKaiNET 0.57.0**.

**Qwen on the compiled IREE KV path.** `QwenExportHarness` / `QwenExportCli` trace Qwen2 and Qwen3
checkpoints to the three graphs of the `qwen-kv-v1` contract (catalog-prefix prefill, one-call chunk,
dynamic-cache decode step), `IreeKvSession` runs grouped-query models, and `Qwen25ChatTemplate` renders
Qwen2.5's official tool-calling template. Qwen3-0.6B compiles for Vulkan and arm32 and reproduces
llama.cpp's greedy tokens exactly on host IREE; it has not been run on a device yet, and Qwen2.5 export
waits on an engine fix for attention bias. **The Moonshine v2 stream** caps its final re-decode at the
model card's length budget, so short commands finish sooner.

**Engine 0.57.0 and Kotlin 2.4.20.** The engine fixes the StableHLO export of an explicit attention mask
under grouped-query attention and onto a dynamic key length, the shape of a chunked-prefill graph over a
KV cache. **The FunctionGemma export contract now carries the embedding geometry** (`vocabSize`,
`nHeads`, `hiddenSize`, `slidingWindow`) into `manifest.json`, read from the checkpoint, so a fine-tune
with added special tokens exports and runs without its new tokens silently collapsing to token 0.

It builds on **0.56.2**, which released `IreeMoonshineStream`, the streaming Moonshine v2 runtime on the
device; **0.56.1**, which released `MoonshineV2ExportCli` (a Hugging Face snapshot in, the five
StableHLO graphs and two host-side tables out, no Python, nothing to configure per language or size), and
**0.56.0**, which put the two version lines back in lock-step with **SKaiNET 0.56.0** (the engine skipped
0.55.0 to realign them).
Expand Down Expand Up @@ -256,7 +262,7 @@ The recommended way to consume is via the BOM. It pins every published `skainet-

```kotlin
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))

// Versions resolved from the BOM:
implementation("sk.ainet.transformers:skainet-transformers-core")
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -32,7 +32,7 @@ From your own build, resolve the published module and run its main class:
----
val exportTool by configurations.creating
dependencies {
exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
exportTool(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))
exportTool("sk.ainet.transformers:skainet-transformers-inference-moonshine")
}
tasks.register<JavaExec>("exportMoonshine") {
Expand Down
2 changes: 1 addition & 1 deletion docs/modules/ROOT/pages/reference/moonshine-encoder.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -24,7 +24,7 @@ a DSL decoder is future work.
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))
implementation("sk.ainet.transformers:skainet-transformers-inference-moonshine")
}
----
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -55,7 +55,7 @@ dependencies {
// self-registers it on ART at process start — nothing to call.
runtimeOnly("sk.ainet.core:skainet-backend-jni-cpu")

implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))
implementation("sk.ainet.transformers:skainet-transformers-core")
implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-inference-llama")
Expand Down
4 changes: 2 additions & 2 deletions docs/modules/ROOT/pages/tutorials/getting-started-java.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -25,7 +25,7 @@ In your `build.gradle.kts`:
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))

implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-agent")
Expand All @@ -41,7 +41,7 @@ Or in Maven (Maven needs the `-jvm` classifier suffix on platform artifacts):
<dependency>
<groupId>sk.ainet.transformers</groupId>
<artifactId>skainet-transformers-bom</artifactId>
<version>0.56.2</version>
<version>0.57.0</version>
<type>pom</type>
<scope>import</scope>
</dependency>
Expand Down
Original file line number Diff line number Diff line change
Expand Up @@ -34,7 +34,7 @@ encoder output to the advertised dimensionality. The runtime applies it automati
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))
implementation("sk.ainet.transformers:skainet-transformers-providers")
}
----
Expand Down
2 changes: 1 addition & 1 deletion docs/modules/ROOT/pages/tutorials/llama3-tool-calling.adoc
Original file line number Diff line number Diff line change
Expand Up @@ -52,7 +52,7 @@ The pieces you need live in three modules:
[source,kotlin]
----
dependencies {
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.56.2"))
implementation(platform("sk.ainet.transformers:skainet-transformers-bom:0.57.0"))

implementation("sk.ainet.transformers:skainet-transformers-runtime-kllama")
implementation("sk.ainet.transformers:skainet-transformers-agent")
Expand Down
2 changes: 1 addition & 1 deletion gradle.properties
Original file line number Diff line number Diff line change
@@ -1,5 +1,5 @@
GROUP=sk.ainet.transformers
VERSION_NAME=0.56.2
VERSION_NAME=0.57.0

POM_DESCRIPTION=SKaiNET-transformers

Expand Down
Loading