Skip to content

Repository files navigation

mumblr

Think out loud, get a prompt. Speech to markdown in the folder you start it from, with edits by voice through your local Claude Code.

ci release platform stack

Releases · CI · Issues

The mumblr window: dictation buffer on the left, command log on the right.

mumblr-<version>-win-Setup.exe installs per user and puts mumblr on your PATH. The portable zip does not: unzip it and add the folder to PATH yourself.

Bring two paid services of your own: an ElevenLabs key, billed per minute of audio, and Claude Code on your PATH if you want the command channel. mumblr is free and talks to nothing besides those two.


A voice recorder built around one loop: thinking out loud while you develop, and turning that into text a coding agent can use. It is deliberately narrow - one markdown file in the folder you start it from, spoken edits handed to your local Claude Code, and the whole buffer on your clipboard the moment you ask for it. Nothing stops you dictating a shopping list with it; nothing is optimized for that either. The roadmap follows the development loop: prompts, notes, commit messages, the things you say to an agent.

mumblr .

That creates dictated-<timestamp>.md in the current folder and opens a window. Next to it go dictated-<timestamp>.wav, the audio, and dictated-<timestamp>.raw.md, what speech-to-text produced (after the dictionary pass) and nothing an LLM wrote. Talk, press stop, record again, run a command over the result - the files stay on disk for any Claude Code session to read by path, and Copy puts the whole buffer on your clipboard when you are actually done.

The two channels

Channel 1 - content. Your speech becomes the text in the buffer. Two interchangeable backends, both built, switchable in the UI:

Mode What it does Keyterms
Realtime (default) Scribe v2 Realtime over a websocket. Committed segments append as you speak, partials only show in the preview line and never enter the buffer. 50, 20 chars
Batch Scribe v2, one POST when you stop. Slower to first text, highest accuracy. 1000, 50 chars

no_verbatim is on, so filler words and false starts are dropped inside the model. The language is auto-detected by default, which handles German with English technical terms out of the box; when a short take in another language throws the detection off, the LANG picker in the toolbar pins it for the next recording and for spoken commands, and stt.languages in the config is the list it offers. Client side there is a deterministic dictionary pass with no LLM involved (clod code -> Claude Code).

Channel 2 - commands. Hold the command key (or the Hold to edit button), say what to change, let go. The clip goes through batch STT, and the resulting command plus the absolute file path go to your locally installed claude -p, which edits the file with its own Read/Edit tools. Typical commands: delete the last sentence, replace X with Y, clean this up, turn this into a prompt. Expect 15-30 s with Opus at high effort. Nothing from this channel ever lands in the content file - it goes to the command log panel with a one-line summary of what changed, written in the language you dictated in, and every call is snapshotted so you can revert it. Revert is one step at a time; Raw puts the dictation back exactly as it was transcribed, whatever the commands did since, and is itself revertible. The raw file is read-only on disk between appends, so a command cannot edit it even though Claude has edit rights in the folder.

Commands you say word for word every day belong on a button instead. Every markdown file in %APPDATA%\mumblr\prompts becomes a button above the log; clicking one skips the microphone and the STT round trip entirely and takes the identical path from there - snapshot, claude -p, reload, revert. Two ship:

  • Grammar fixes grammar, sentence structure and word order and changes nothing else: not the content, and not its language. German dictation with English terms comes back as German dictation with English terms.
  • Prompt is the "get a prompt" in the tagline. It shapes the dictation into something you can hand to a coding agent - the ask first, then the context you gave, then the constraints - and ends with an Open questions section holding every gap and contradiction as a question. It adds nothing you did not say: the agent receiving the prompt can ask, you can answer, and a shaping that fills the gaps for you would be inventing requirements about a codebase it has never seen. Raw is one click away if it went too far.

States

Exactly one writer at a time:

State Editor Writer
Idle free you
Recording locked STT, appending at the marker set when recording started
Commanding locked claude -p, buffer reloaded from disk afterwards

Starting a command while recording pauses channel 1 and resumes it when the command is done.

While recording, the window title says so and the taskbar button flashes whenever mumblr is not the window in front. A recording forgotten behind the IDE keeps the microphone open and the meter running; the flash is for the one state only you can end.

Setup

The ElevenLabs key comes from the environment only, never from a config file or the repo:

setx ELEVENLABS_API_KEY "your-key"    # XI_API_KEY also works

claude must be on your PATH for channel 2.

The installer adds mumblr to your user PATH and removes it again on uninstall, so mumblr . works from any folder. A PATH change only reaches new terminals - the one you installed from still will not find it. For the portable build, add the folder you unzipped to PATH by hand:

$dir = "C:\where\you\unzipped"
[Environment]::SetEnvironmentVariable('Path', "$([Environment]::GetEnvironmentVariable('Path','User'));$dir", 'User')

Everything else lives in %APPDATA%\mumblr\config.json (the Config button opens it). One file, shared by every window: mumblr . is meant to be run per repo folder, and a setting changed in one window reaches the others by itself. A change that arrives during a recording or a running command waits for it to end rather than swapping the microphone underneath it.

Setting Meaning
microphoneDeviceId The chosen capture endpoint. mumblr never falls back to the Windows default; if the device is gone it shows the picker.
sttMode Realtime or Batch
keyterms Priority ordered. The head of the list survives the realtime limit of 50. A term carrying < > { } [ ] \ or more than five words is dropped - ElevenLabs refuses the whole request over one bad term. Past 100 terms every request is billed as at least 20 seconds, and keyterms carry a 20% surcharge.
dictionary Literal replacements applied to committed text
hotkeys enabled (the status bar toggle, off on a fresh install), toggleRecording, copy, revertCommand, commandHoldKey
claude model, effort, headerPrompt, allowed/disallowed tools, timeout
stt Model ids, noVerbatim, languageCode (what the LANG picker chose, unset for auto), languages (what it offers), base URL, VAD silence threshold, keytermsEncoding

Two settings are deliberately not in the file. ELEVENLABS_API_KEY (or XI_API_KEY) carries the transcription key, and MUMBLR_GITHUB_TOKEN lets the updater read the release feed of a private repository. Both come from the environment only, never from config and never from the repo.

Default hotkeys: Ctrl+Alt+Space record, hold Ctrl+Alt+D for a command, Ctrl+Alt+C copy, Ctrl+Alt+Z revert. They work while your IDE or terminal has focus - which is the point, and also the risk: a chord that collides with a game or another tool starts a recording you did not want. So a fresh install starts with them off: the buttons do everything, and the hotkeys toggle in the status bar turns the chords on once you have decided you want them. The same toggle turns them off again with one click, unregistering the chords and removing the keyboard hook. The state is saved either way. An upgrade keeps whatever you had.

Your prompts

The command buttons are files. %APPDATA%\mumblr\prompts\*.md, one prompt each, written on the first run and yours from then on - add one, edit one, delete one, and the buttons follow without a restart. The Prompts button in the toolbar opens the folder. Frontmatter names the button and places it; everything below it is what goes to Claude:

---
label: Shorter
order: 30
---

Halve the length without losing a single point. Keep the author's words and language.

Both keys are optional: a file with no frontmatter is a prompt named after itself, and one with no order sorts after every file that has one. Anything else between the fences - a comment, a key mumblr does not know - is ignored, and a block that names neither label nor order is not read as frontmatter at all, so a prompt may open with a horizontal rule or a line like Rule: keep it short and keep it. A file that holds no prompt, or that is not UTF-8, gets no button and the window says which one and why.

The directory is the only place prompts come from, and it is not configurable. A prompt goes to claude -p with permission to read and edit the dictation file, so a prompt directory inside a repository would let a repo you cloned run its own instructions over what you dictate.

Deleting a file keeps it deleted. Deleting the whole folder is how you ask for the shipped two back: press Prompts afterwards, or restart. What mumblr goes by is the .seeded file it writes in there once - delete that alone and it writes back whichever of the shipped prompts are missing. One it wrote and you then edited counts as missing: it comes back beside your copy rather than over it, so you get two buttons and delete the one you do not want.

Preview builds

Releases come in two flavours, and the download you install decides which one you get. A normal release (v0.2.1) is stable; one tagged v0.2.1-beta.1 is a preview, marked as a pre-release on the releases page, with -win-beta- in its filename. The two update independently: an installed build only ever offers updates from its own release channel, so a preview never arrives on a stable install and a stable release never quietly replaces a preview.

There is no switch in the window. Install a preview build to get on previews, install a stable one to get back - both are per-user installs over the same app, and your config.json is untouched either way.

What it deliberately is not

  • No wake word and no hands-free control - you press something to record.
  • No cursor injection into other applications. The output is a file and the clipboard.
  • No diarization and no long recordings. It is built for a few minutes of thinking out loud.
  • No prompt library with sharing, versions or variables. A prompt is a markdown file you own.
  • No cloud LLM inside the app. The only model that touches your text is the Claude Code you installed yourself.
  • No editing while a recording runs - one writer at a time, by design. See the state table above.

Compliance

Audio is stored by ElevenLabs on standard tiers, not just processed. Accepted risk - so do not dictate anything that could not go into an external prompt: no credentials, no customer names, no ticket internals. LLM processing happens only through the locally installed Claude.

Two hosts are contacted, and no others. ElevenLabs receives the audio. GitHub is asked for the release feed at startup and when you click the version button - that request carries no dictation, only a version check. There is no telemetry.

Building

dotnet test
dotnet publish src/Mumblr.App/Mumblr.App.csproj -c Release -r win-x64 --self-contained -o publish

dotnet test never touches the network. The handful of tests that do talk to ElevenLabs are armed separately, because they cost money on every run:

MUMBLR_LIVE_TESTS=1 dotnet test --filter FullyQualifiedName~LiveElevenLabs

They exist because the rest of the suite can only assert what mumblr sends, not what ElevenLabs accepts - which is how the keyterm encoding shipped broken past a green build.

Project What it holds
src/Mumblr.Core State machine, STT interface and backends, audio pipeline, config, claude -p invocation. Platform independent and unit tested.
src/Mumblr.App Avalonia UI with AvaloniaEdit, WASAPI capture, Win32 hotkeys, Velopack.
test/Mumblr.Core.Tests Core unit tests
test/Mumblr.App.Tests Headless Avalonia tests over the whole view model, both channels

The STT interface exists so a local backend (Whisper, Parakeet, CPU, batch) can be dropped in later without touching the UI.

Releases are built by GitHub Actions on Windows and packaged with Velopack: tag vX.Y.Z and push.

About

Think out loud, get a prompt. Speech to markdown in the folder you start it from, with edits by voice through your local Claude Code.

Topics

Resources

Security policy

Stars

1 star

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages