Local-first audio & video transcription for Apple Silicon Macs.
NVIDIA's Parakeet ASR running entirely on-device (CoreML / Neural Engine), with speaker diarization, a queue-driven GUI, and a scriptable CLI — one engine behind both.
- Fully local: audio never leaves your Mac, and everything works offline after the first model download.
- Fast: the default speech-swift model transcribes at ~45–95× real-time on Apple Silicon; other model speeds vary.
- Lean: the default model used ~1.7 GB peak memory in a 2-hour test; FluidAudio long-file memory is still being validated.
- Speaker-aware: name up to 4 attendees and every line comes out labeled with the right speaker.
- App + CLI: choose a model per job in the queue app or script the CLI with JSON output; both share the same pipeline.
- Formats: clean Markdown with optional timestamps, or ready-to-use SRT subtitles with speaker prefixes.
- Download the latest
Conure-<version>.dmgfrom Releases. - Open the DMG and drag Conure to Applications.
- First launch: right-click the app → Open (it's unsigned — see below), and the required models (~240 MB) download automatically.
Unsigned app note: Conure ships without a Developer ID. If macOS blocks it, right-click → Open, or clear the quarantine flag manually:
xattr -cr /Applications/Conure.app
Requirements: macOS 15+ on Apple Silicon (M1–M4). ~860 MB of model downloads on first use (auto-managed).
App — launch Conure, hit +, drop in files (optionally name up to 4 speakers), and watch the queue run.
CLI — install the bundled engine on your PATH from Settings, or use it straight from a build:
# Plain transcript (Markdown, no timestamps) next to the input file
conure transcribe meeting.mp4
# Speaker-labeled Markdown with timestamps and attendee names
conure transcribe meeting.mp4 --speakers "Alice,Bob" --timed
# SRT subtitles with speaker prefixes
conure transcribe recording.m4a --speakers "Alice,Bob" --format srt
# Multiple files (processed sequentially), custom output folder
conure transcribe *.mp4 --speakers "Alice,Bob,Carol" --output ~/transcripts
# Existing transcripts are never silently overwritten — pick a policy:
conure transcribe meeting.mp4 --replace # overwrite existing output
conure transcribe meeting.mp4 --unique # write "meeting 2.md" instead
# Progress as JSON lines (for tooling); --pretty for humans
conure transcribe talk.wav --pretty
# Punctuated English or multilingual transcription (models download on first use)
conure transcribe meeting.mp4 --model parakeet-unified-en
conure transcribe french.wav --model parakeet-v3-multilingual --language fr| Flag | Default | Description |
|---|---|---|
--model |
parakeet |
Model id or HuggingFace repo |
--language |
auto | Optional script hint for multilingual v3 (e.g. fr); Japanese auto-detects without a hint |
--speakers |
— | Comma-separated attendee names (max 4); enables diarization |
--format |
md |
md or srt |
--timed / --no-timed |
no | Timestamps in Markdown output ([H:MM:SS] per line) |
--output |
input folder | Output directory |
--replace |
off | Overwrite existing output files |
--unique |
off | Write to a numbered filename (name 2.md) when the output already exists |
--pretty |
JSON lines | Human-readable progress |
Re-running a transcription never silently overwrites an existing file: by default the CLI stops with an error if the output path is already taken, including when two inputs in the same invocation share a basename and output folder. Pass --replace to overwrite or --unique to write a numbered copy (meeting 2.md) that keeps the correct extension. The queue app performs the same check when jobs are added and asks whether to replace, keep both, or cancel. Output files are always written atomically after the destination is resolved.
conure models list # registry, download status, sizes (also --json)
conure models download parakeet # pre-download (~611 MB)
conure models download parakeet-unified-en
conure models download parakeet-v3-multilingual
conure models remove parakeet # free disk spaceModels live in ~/Library/Application Support/Conure/models/ (override with CONURE_MODELS_DIR). FluidAudio weights are stored in its FluidAudio/ subfolder, downloaded through the SDK and removable with conure models remove. Missing models download automatically on first use. models list --json includes the owning engine.
| Id | Purpose | Size | Removable |
|---|---|---|---|
parakeet |
ASR — Parakeet TDT 0.6B v2 English, CoreML INT8, Neural Engine (default) | ~611 MB | Yes |
parakeet-unified-en |
FluidAudio Unified EN, punctuated and capitalized | ~590 MB | Yes |
parakeet-v3-multilingual |
FluidAudio Parakeet TDT v3, 25 European languages + Japanese | ~465 MB | Yes |
sortformer |
Speaker diarization (≤4 speakers), used with --speakers |
~239 MB | No (required) |
silero |
Voice activity detection, used otherwise | ~1 MB | No (required) |
The original parakeet remains the default. FluidAudio provides the other two ASR models; Sortformer and Silero remain on speech-swift for every model.
- Decode — AVFoundation extracts mono 16 kHz PCM (mp4, mov, m4a, mp3, wav, aac, aiff)
- Segment — with
--speakers: Sortformer diarization produces speaker turns; otherwise Silero VAD finds utterances. Same-speaker turns merge into ≤25 s chunks - Transcribe — each chunk through the selected ASR engine; chunk bounds become line timestamps, so a line can never blend two speakers
- Write — Markdown or SRT next to the input (or
--outputdir)
The GUI drives the exact CLI as a subprocess (JSON progress on stdout) — one engine everywhere. Internals and repo layout are documented in AGENTS.md.
Pull requests are welcome.
- Fork the repo and create a branch using
<type>/<issue>-<slug>(e.g.feat/42-add-export). - Open a pull request against
maindescribing the change.
Coding agents should follow AGENTS.md for project structure, conventions, and setup.
- speech-swift (Apache-2.0) © soniqo — the default inference engine, vendored with a long-file fix at
vendor/speech-swift/ - FluidAudio (Apache-2.0) © FluidInference — Unified EN and multilingual CoreML inference
- Parakeet weights (see each model card for license and attribution; Unified EN: CC-BY-4.0) © NVIDIA; CoreML conversions © their respective publishers
- Sortformer © NVIDIA
- Silero VAD (MIT)
Conure code is released under the MIT License.