VibeASR.cpp is the official inference runtime for VibeVoice-ASR-BitNet β enabling real-time multilingual speech recognition on CPU through heterogeneous quantization (I8_S for VAE + I2_S for LM).
To enable efficient edge CPU deployment, we replace the original Qwen2.5-7B language model with Qwen2.5-1.5B, achieving only modest accuracy degradation (1β4% absolute WER increase) while reducing the total model size from 4.62 GB to 1.58 GB. Combined with custom SIMD kernels and operator fusion in the ggml framework, VibeVoice-ASR-BitNet achieves 1.6β2.3Γ faster inference than Whisper.cpp at comparable model sizes, with real-time capability (RTF < 1) on low-resource CPUs.
π Tech Report Β |Β π€ Models Β |Β π GeneralAI
| Component | VibeVoice-ASR-1.5B (FP16) | VibeVoice-ASR-BitNet | Compression |
|---|---|---|---|
| VAE Tokenizer | 1.31 GB | 0.65 GB | 2.0Γ |
| LM Decoder | 3.32 GB | 0.92 GB | 3.6Γ |
| Total | 4.62 GB | 1.58 GB | 2.9Γ |
| Threads | RTF | vs. Whisper.cpp |
|---|---|---|
| 1 | 1.98 | 2.28Γ |
| 2 | 1.08 | 2.12Γ |
| 3 | 0.77 | 1.86Γ |
| 4 | 0.63 | 1.86Γ |
| 6 | 0.49 | 1.71Γ |
| 8 | 0.42 | 1.55Γ |
Benchmarked on AMD EPYC 7V13 (AVX2+FMA) with 20s audio input. Bold = RTF < 1 (real-time).
| Benchmark | VibeVoice-ASR-7B (FP16) | VibeVoice-ASR-BitNet | Parakeet | Whisper | SenseVoice | FunASR |
|---|---|---|---|---|---|---|
| MLC-EN | 7.82 | 8.25 | 8.40 | 13.57 | 12.39 | 11.36 |
| MLC-FR | 16.03 | 17.41 | β | β | β | β |
| MLC-IT | 15.67 | 17.23 | β | β | β | β |
| MLC-KO | 9.83 | 11.15 | β | β | β | β |
| MLC-PT | 22.41 | 24.87 | β | β | β | β |
| MLC-VI | 20.15 | 22.38 | β | β | β | β |
| AISHELL4 | 19.83 | 27.45 | β | β | 22.52 | 20.41 |
| AMI-ihm | 17.42 | 21.36 | 21.92 | 27.07 | 30.81 | 32.07 |
| AMI-sdm | 24.18 | 25.87 | 26.33 | 36.92 | 48.11 | 40.17 |
| AliMeeting | 36.21 | 40.58 | β | β | 38.75 | 39.27 |
| Fleurs-en | 4.73 | 5.21 | 4.09 | 3.99 | 6.84 | 4.93 |
| Fleurs-zh | 7.92 | 8.35 | β | β | 5.56 | 7.00 |
| Libri-clean | 2.17 | 2.41 | 1.49 | 1.98 | 2.78 | 1.58 |
| Libri-other | 5.84 | 6.27 | 3.13 | 3.60 | 6.81 | 4.01 |
| VoxPopuli | 4.92 | 5.18 | 5.26 | 7.19 | 8.63 | 6.46 |
- Python β₯ 3.9, CMake β₯ 3.14, GCC/Clang with C++11 support
- ~2 GB disk space (code + quantized models)
git clone --recursive https://github.com/microsoft/VibeASR.cpp.git
cd VibeASR.cpp
pip install -r requirements.txt
python setup_env.pygit clone --recursive https://github.com/microsoft/VibeASR.cpp.git
cd VibeASR.cpp
# Build
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j$(nproc)
# Download pre-quantized models
pip install huggingface_hub
huggingface-cli download microsoft/VibeVoice-ASR-BitNet --local-dir models/vibeasr./build/bin/asr_infer \
--vae-model models/vibeasr/vibeasr-vae-encoder-i8_s.gguf \
--lm-model models/vibeasr/vibeasr-lm-i2_s-embed-q6_k.gguf \
--audio input.wav -t 4pip install gradio soundfile numpy
python demo/gradio_asr_demo.py --port 7860 \
--vae-model models/vibeasr/vibeasr-vae-encoder-i8_s.gguf \
--lm-model models/vibeasr/vibeasr-lm-i2_s-embed-q6_k.ggufFor most users, downloading pre-quantized models from HuggingFace is recommended. To convert from SafeTensors yourself:
# LM (BitNet) β handles weight preprocessing and config flattening automatically
python utils/convert_lm_to_gguf.py <safetensors-dir>
# VAE Tokenizer
python utils/convert_vae_to_gguf.py <safetensors-dir># VAE Tokenizer: F32 β I8_S
./build/bin/llama-quantize \
<safetensors-dir>/vibeasr-vae-encoder-f32.gguf \
<safetensors-dir>/vibeasr-vae-encoder-i8_s.gguf \
I8_S 1 1
# LM: F32 β I2_S (with Q6_K embeddings)
./build/bin/llama-quantize --token-embedding-type Q6_K \
<safetensors-dir>/vibeasr-lm-f32.gguf \
<safetensors-dir>/vibeasr-lm-i2_s-embed-q6_k.gguf \
I2_S 1 1@article{xu2025vibeasrbitnet,
title={VibeVoice-ASR-BitNet Technical Report},
author={Xu, Songchen and Song, Ting and Huang, Shaohan and Peng, Zhiliang and Xia, Yan and Tu, Yujie and Huang, Xin and Yu, Jianwei and Dong, Li and Wei, Furu},
journal={arXiv preprint arXiv:2607.21075},
year={2025}
}This project is licensed under the MIT License.
