Stop talking. The words
are already at your cursor.

Talking is the fast part. Typing is the second pass — and so is cleaning up what a transcriber hands back. Hold a key, say it, and finished text lands in whatever field the cursor is already in: a chat box, an email, a commit message, a prompt.

Download for macOS See how it works

On your phone? Ask for a seat, or see what the tidying does →

Right ⌘Dictate Right ⌥Translate fnAsk EscCancel
your cursor is here

This demo is drawn in code, not recorded — the capsule here is the one you get in the app.

Subtract · One step fewer

One sentence, two ways to finish it.

Typing 0:00
Type, reread, delete, rearrange…
Speaking 0:00
your cursor is here
✓ Inserted. On to the next thing.

A few dozen inputs a day, and a minute or two saved on each one. That is the whole argument this page is making.

Four Things · What the keys do

Four keys, four things.

01
Dictate

Say it, get it written. The “um”, the “you know”, the false start get cleaned out on the way; what lands is a sentence you would have typed.

Right ⌘
02
Translate

Speak English, Chinese lands; speak Chinese, English lands. The direction is worked out for you, so a message to a colleague in another language is one held key.

Right ⌥
03
Ask

Say the question and what lands is the answer, not your question. A flag you half-remember, a conversion, a spelling — without leaving the window.

fn
04
Edit by voice

Select a paragraph and say how it should change — what lands back is the changed version. Being able to revise by mouth is what closes the loop.

Right ⌃
DictateRight ⌘
your cursor is here
TranslateRight ⌥
your cursor is here
Askfn
your cursor is here
Vibe Coding · Talking to the tools

Telling Claude Code, Cursor or Codex what you want
was always a spoken job.

You are not writing the code character by character any more — you are explaining what you want. “Make this function async, then put a retry around it.” Three seconds to say; thirty to type. Murmur has no plugin, no integration and nothing to set up — it is system-level dictation, and it lands at any cursor, including the prompt in your terminal that is waiting for input.

A long prompt stops costing you anything.The constraint you normally leave out because typing it was tedious — say it instead. A spoken paragraph takes about as long as a spoken line.
Get a name wrong once, correct it once, never again.Fix the library or API name it misheard and it goes into your dictionary; next time it is heard right — conservative on purpose: it would rather miss one than invent one.
Your standing prompt, saved as a phrase.Say the trigger word and the whole block lands — the rules you re-explain every single day, typed out once and only once.
Check an API name without switching windows.Hold Ask, say the question, and what lands is the answer, not your question — hands never leave the editor.
Select some code and say how to change it.Edit by voice puts the changed version back where it was — the thought never has to be typed first.
In an editor it leaves your identifiers alone.Tone per app can tell an editor from a chat window, and does not apply the same rules to both.
Engines · What is behind the switch

The recognition engine is swappable.
This is the shortlist, after we did the testing.

We put the recognition models we could get our hands on through one internal evaluation and kept three. None of them wins everywhere — how you talk is what decides it — so they are split by situation: strengths and weaknesses printed on the front, one menu to switch, no terminology to learn first.

Standard

1× quota
The default for everyday speech
Everyday use●●●●●
Chinese-heavy●●●●
English-heavy●●●●●
  • Live captions; final text the moment you let go
  • Steadiest when languages are mixed mid-sentence (our own testing)
  • Lightest on quota
  • Mid-pack on public streaming benchmarks for long English sentences

Chinese-first

1× quota
Mostly Chinese, regional accents
In closed beta · rolling out
Everyday use●●●●
Chinese-heavy●●●●●
English-heavy●●●●●
  • Best published results on Chinese and on mixed speech
  • Holds up on regional accents (Cantonese, Sichuanese, Wu)
  • Live captions; final text the moment you let go
  • No proper-noun biasing in the live captions yet (your dictionary still fixes the final text)

Precision

2× quota
Mostly English, dense with names
In closed beta · rolling out
Everyday use●●●●●
Chinese-heavy●●●●●
English-heavy●●●●●
  • Front rank on public English and mixed-speech benchmarks
  • Steadiest on heavy accents and on proper nouns
  • No live captions (you wait for the whole passage)
  • Chinese output occasionally mixes simplified and traditional
  • Heaviest on quota

Standard is the tier open to everyone today; the other two are opening up through the closed beta. “2× quota” means exactly that: on that tier a minute of long recording takes two minutes out of the monthly pool, and dictation draws from your weekly allowance at twice the rate. The multiplier is printed next to the tier; there is no hidden conversion. Lose the network or run out of allowance and it falls back to on-device recognition by itself: slower, weaker when you mix languages, but never a brick, and never a lost sentence.

Local · Where your data is

Our server knows who you are.
It never learns what you said.

Accounts, sign-in, subscriptions — that is the entire job our server does. Your recordings, transcripts, history and dictionary are never uploaded; they exist on your Mac only. For cloud recognition the audio goes straight from your Mac to the provider you picked, not through us. Choose on-device recognition and not one byte leaves this computer.

Your Mac audio, history, dictionary ✓ Stored locally Our server knows only who you are Your provider one tier, one destination

You speak — the audio leaves from your Mac and nowhere else

~/Library/Application Support/Murmur/
The recordings, and every word you have said, are in that folder. Open it whenever you like. Delete it and it is actually deleted.

Promises · Things we will put in writing

Four of them. Written down means kept.

Never lose what you said

Every dictation keeps the audio and a history entry. Cloud down, recognition wrong? Open the history and run it again on another engine. No path through this app makes a sentence you said disappear.

Failure gets said out loud

If the translation did not happen, it tells you it did not happen. If the microphone heard nothing, it says so on the spot. It will not quietly push some fallback mess into the window you are typing in.

Where the audio goes is on the surface

On-device: not one byte leaves this computer. Cloud: which provider, and what got sent, is written in Settings in plain words — not on page eight of the terms.

The numbers are measured

The “release the key to text on screen” latency is a measured median, not a number a copywriter picked. Each engine's weak spots are printed next to its strong ones — switch if it does not suit you; your ears decide.

Lecture · Long recordings

Leave it running; captions float over whatever is on screen.

Long recording is a separate pipeline. The caption window floats above any full-screen window — the talk keeps playing, the captions keep up; turn on the second column if you want a translation beside them. Captions are an expendable preview: if they drop, the recording does not. The full transcript lands in your history when you stop, with the speakers told apart.

Live captionspreview · the full transcript still lands in history

Drawn in code, not recorded. Free tier: 30 minutes a month. Pro: no limit on length.

Details · Small things that pull their weight

Not in the headlines, but working every day.

The dictionary grows itself. Proper nouns you have corrected go into the dictionary and are heard right next time — conservative on purpose: it would rather miss one than invent one.
Quick phrases. Say the trigger word and a whole block lands — an address, payment details, a reply you send twice a week. That path never touches a large model.
Tone per app. It lands in a chat app as speech, in an email as prose, and in a code editor without touching your identifiers.
Said it wrong? Fix it in the history. Every dictation keeps its original text and its audio; change one word and that becomes the reference from then on.
Under-your-breath mode. Say it almost silently in a library; the capture end amplifies and recognises it as usual.
“Remind me in half an hour.” Say it in passing and the reminder is set; it turns up on time — handled on-device, so it costs no allowance and works offline.
Names on your screen help it hear right. At the moment you start recording it reads the visible text of the front window once to pull out proper nouns (no screenshot; screen words never enter your dictionary) — the details are in “Where your data lives”.
Questions · The ones you are about to ask
Where does my audio get sent?

Your recordings, transcripts, history and dictionary all live on your Mac; our server only handles accounts and subscriptions. For cloud recognition the audio connects straight to that one recognition service and does not pass through our server — that holds while you are on the free allowance we pay for, too. One exception, stated plainly: on the free allowance the text for the cleanup step goes through our server before it reaches the model provider. Pick on-device recognition and nothing leaves the computer at all. Proper nouns from your dictionary are sent along with the request, so that they get heard correctly — all of this is written in Settings, not buried in terms.

What happens when I am offline?

Offline, or out of allowance, it falls back to on-device recognition by itself — slower, weaker when you mix languages, but your words are not lost. Once you are back, any entry in the history can be run again on a cloud tier.

Which languages does it handle?

English, Chinese, and the two of them mixed inside one sentence — that last case is the one it was built to take seriously. Chinese with a regional accent (Cantonese, Sichuanese, Wu) has a tier of its own, in closed beta and rolling out.

Do I have to bring my own API key?

No. Cloud allowance is invite-only for now: get a seat, sign in, and the free allowance is there — it is a cloud bill we pay for you, and when it runs out the app returns to on-device recognition rather than stopping to ask you for money.

Try saying the next one out loud.

Download for macOS

On your phone? Ask for a seat, or see what the tidying does →

Requires macOS 26 or later on an Apple Silicon (M-series) Mac · Free to start · Cloud allowance by invite

On an Intel Mac, or still on macOS 25 or earlier? It won't install yet — leave your email at foraiandfocus+murmur@gmail.com and we'll tell you the moment it's supported.

The download is free and on-device recognition works without an account. The cloud allowance is invite-only for now — email foraiandfocus+murmur@gmail.com and tell us in one line what you'd use it for.

Unzip it and drag Murmur into Applications. On first run it asks for the microphone and for accessibility — the first to hear you, the second to put the words at your cursor.