A Home Assistant custom integration providing local, on-device speech-to-text via the Cortex STT Server -- a multi-model server running Whisper, NVIDIA Parakeet, SenseVoice, Qwen3-ASR, and more on a single GGUF runtime.
Audio ──► Cortex STT Server ──► Transcribed Text
│
├── Whisper (multilingual)
├── Parakeet (English, low latency)
├── SenseVoice (Asian languages)
└── Qwen3-ASR, Canary, Moonshine, ...
Each downloaded model on the server becomes its own STT entity in Home Assistant, so voice pipelines can pick the right model per language or per use case.
Wrong characters for Chinese device names? Pair with STT Corrector -- it normalizes language, applies custom replacements, and pinyin-matches transcripts against your HA areas and devices.
- Multi-model STT -- Whisper, Parakeet, SenseVoice, Qwen3-ASR, and more served by the Cortex STT Server
- WebSocket streaming -- audio is fed to the server while you are still speaking, so decoding overlaps capture; falls back to a single POST automatically if the stream cannot be opened
- Per-model entities -- one STT entity, one
model loadedbinary sensor, and eleven diagnostic sensors per downloaded model - Automatic model discovery -- discovers everything on the server at setup, then adds/removes entities live as models are downloaded or deleted server-side (no reload needed)
- Supervisor auto-discovery -- when the Cortex STT app is installed, the integration is offered automatically with the URL and API key pre-filled (no manual setup needed)
- Capture-device attribution -- each transcription is tagged with the Assist satellite that recorded it, so the server's history can compare transcription quality per microphone (works through STT Corrector too)
- Runtime statistics -- diagnostic sensors track request counts, inference duration, audio duration, and real-time factor
Prerequisites: Home Assistant 2026.3.0+ and a running Cortex STT Server instance (available as a Home Assistant app).
Click the button above, or manually: HACS > three-dot menu > Custom repositories > add https://github.com/hass-cortex/cortex-stt (Integration) > install > restart HA.
Manual installation
Copy custom_components/cortex_stt/ to your HA config/custom_components/ directory, then restart.
Install and start the Cortex STT Server, then download at least one model. Note the server's URL (e.g. http://homeassistant.local:8769) and API key.
When the Cortex STT Server runs as a Home Assistant app, setup is automatic — Home Assistant Supervisor discovers the app and shows a "Cortex STT discovered" card under Settings > Devices & Services. Click Configure and the URL and API key are filled in for you; just confirm to create the entry.
For non-app servers (or if auto-discovery is disabled), add the integration manually:
Click the button above, or manually: Settings > Devices & Services > Add Integration > search "Cortex STT". Enter the Server URL and API Key. The integration validates connectivity and credentials before completing setup, then creates one device per downloaded model.
Select or create a voice pipeline, then set Speech-to-text to the Cortex STT entity for the model you want to use. If you have several downloaded models, you can create one pipeline per model and switch between them per use case.
Open the integration page and click Configure to adjust:
| Option | Default | Range | Description |
|---|---|---|---|
Polling interval (update_interval) |
30 s |
5 – 3600 | How often the Model loaded binary sensor polls the server's /api/engine endpoint. Lower values detect load changes faster at the cost of more HTTP traffic. |
Changes take effect immediately -- the integration reloads itself when you save.
Settings > Devices & Services > Cortex STT > three-dot menu > Delete > (HACS) remove the repository or (manual) delete custom_components/cortex_stt/ > restart HA.
Per downloaded model, the integration creates:
| Entity | Type | Description |
|---|---|---|
| STT | stt |
Speech-to-text entity for voice pipelines |
| Total requests | sensor |
Total transcription requests (diagnostic) |
| Successful requests | sensor |
Successful transcription count (diagnostic) |
| Failed requests | sensor |
Failed transcription count (diagnostic) |
| Last inference duration | sensor |
Duration of the last transcription (ms) |
| Average inference duration | sensor |
Rolling session average duration (ms) |
| Last audio size | sensor |
Size of the last audio input (bytes) |
| Total audio duration | sensor |
Cumulative audio processed (minutes) |
| Last audio duration | sensor |
Duration of the last audio input (seconds) |
| Transcribed text | sensor |
Last successfully transcribed text |
| Last result | sensor |
Last result status (success / no_speech / api_error) |
| Real-time factor | sensor |
Inference time / audio duration ratio |
| Model loaded | binary_sensor |
Whether the model is loaded in memory |
Several diagnostic sensors are disabled by default to keep the UI tidy -- enable them individually from the device page if you want to graph them.
- Multilingual household -- pair a
zhWhisper pipeline and anenParakeet pipeline, each routed to the matching Cortex STT entity, so every member of the family speaks to Assist in their preferred language. - Low-latency wake-to-action -- use Parakeet for short command pipelines (low real-time factor) and reserve Whisper-Large for longer dictation pipelines that value accuracy over speed.
- Quality monitoring -- dashboard the
Real-time factorandAverage inference durationsensors to spot GPU throttling or model regressions, and chartTotal audio durationagainst your hardware utilisation. - A/B model comparison -- keep two models downloaded on the server, assign each to a different pipeline, and compare the
Last raw textsensors for representative phrases.
Enable debug logging to trace transcription requests and server responses:
# configuration.yaml
logger:
default: info
logs:
custom_components.cortex_stt: debugWhy does my pipeline show "no speech" for short utterances?
The Cortex STT Server requires non-silent audio. Check that your wake-word / VAD is cutting audio cleanly and that the Last audio duration sensor is non-zero. If the server replies successfully but with empty text, the integration reports no_speech (not api_error).
How do I pick between Whisper, Parakeet, SenseVoice, and the rest?
Each model is exposed as its own STT entity -- just assign the one you want to a voice pipeline. Rough guidance: Parakeet for low-latency English, Whisper for multilingual accuracy, SenseVoice for Asian languages; the server's catalog marks recommended models per family. Use the Real-time factor sensor to compare live performance on your hardware, and the server's history (audio level + capture device) to tell microphone problems apart from model problems.
My pipeline uses zh-TW but the server only reports zh. Does that work?
Yes. The integration expands each base language code the server advertises (e.g. zh, en) to a curated list of BCP-47 locale variants (zh-TW, zh-CN, en-US, ...) so Home Assistant's pipeline language matching succeeds.
Transcription sometimes fails with an API error -- what should I check?
- Confirm the Cortex STT Server is reachable from HA (the
Model loadedbinary sensor goesunavailablewhen polling fails) - Inspect the server's logs for the failing request
- Check HA logs with
custom_components.cortex_stt: debugenabled (see Debugging) -- failed requests are logged with the model ID and exception - If the server rejects the API key, the integration triggers a reauth flow automatically
Can I connect to multiple Cortex STT Server instances?
Yes. Add the integration multiple times, once per server URL. Each instance is keyed by a hash of the server URL so duplicates are rejected automatically.
How do I install the latest development version?
After the integration is installed via HACS, switch to the latest main branch using the update.install action:
- Go to Developer Tools > Actions
- Select the
update.installaction - In Target, select the Cortex STT update entity (e.g.
update.cortex_stt_update) - In Version, enter
main(or a specific commit hash) - Click Perform Action, then restart HA
Development versions may contain breaking changes -- revert by running the same action with a release tag.
- Audio format is fixed -- the STT entity accepts 16 kHz / 16-bit / mono PCM WAV only. Voice pipelines already produce this format, but custom integrations sending audio directly must match.
See CONTRIBUTING.md for development setup, testing, and contribution guidelines.
The Cortex STT Server app is built on top of transcribe.cpp -- the single GGUF/ggml runtime powering every model family.