Guides

How to dictate code, symbols, and technical text

Voice typing for developers is finally practical. How to speak identifiers, operators, paths, and numbers — and what a dictation engine must do to keep up.

Developers were promised voice coding for a decade, and mostly got demos. What changed recently isn’t speech recognition — it’s the formatting layer on top. Recognizing “max retries equals three” was always easy. Producing max_retries = 3, and knowing that this sentence wanted code and not the words “max retries equals three,” is the actual problem.

Here’s what works today, and how to get the most out of it.

Where voice already beats typing for developers

Be strategic: the highest-value dictation targets aren’t source files.

  • Commit messages and PR descriptions. You already know what you did; saying it is 3× faster than typing it, and nobody reviews your WPM.
  • Code review comments. The friction of typing makes review comments terse to the point of rudeness. Spoken comments come out complete.
  • Chat with AI coding agents. Prompting an agent is conversation, and the words come faster out loud — especially multi-paragraph context.
  • Docs, ADRs, incident notes. Prose about systems, full of identifiers.

The common thread: technical prose with embedded symbols — exactly what a general transcriber mangles.

The vocabulary of speaking symbols

A good engine accepts the way people naturally read code aloud:

You say You get
“snake case max retries equals three” max_retries = 3
“camel case fetch user profile” fetchUserProfile
“status double equals four oh four” status == 404
“arrow function x fat arrow x times two” x => x * 2
“and and” / “or or” && / ||
“open paren … close paren” (…)
“tilde slash projects slash api” ~/projects/api

The critical property is determinism. If “double equals” produces == on Monday and “= =” on Tuesday, you’ll stop trusting it and proofread everything — at which point dictation has negative value. This is why Meosu’s spoken-form layer is a deterministic engine, not a suggestion from a language model: the same phrase produces the same characters, every time, and the whole behavior is verified against a public 130-case corpus with zero edits required.

Ambiguity is the hard part

The engine also has to know when not to fire. “The period drama we watched” must not become “The . drama we watched.” “I have two options” must not become “I have 2 options” in a sentence where prose style wants the word. Guarding these boundaries is most of the real engineering, and it’s where you’ll feel the difference between engines within an hour of real use.

Practical setup tips

  1. Add your identifiers to a dictionary. Project names, service names, teammates’ names. Thirty seconds of setup removes the most common recognition errors. Meosu supports bulk CSV import.
  2. Dictate into the tool, not into a scratch file. Insertion at the cursor — in your terminal, IDE chat, or PR form — is where the speed lives. Copy-paste round trips eat the gain.
  3. Learn one undo phrase. In Meosu, “scratch that” reverses the last insertion exactly, within 30 seconds. Recovery being cheap is what makes speaking freely feel safe.

Try the test

Read one of your own recent commit messages out loud, symbols and all, into whatever dictation you use. If the result compiles as English and as formatting, keep it. If not — we built Meosu for exactly this, and on-device dictation, which is free, covers a lot of commit messages.

Common questions

Can you really write code by voice?

Dictating prose about code — commit messages, reviews, docs, chat — works extremely well today. Dictating dense syntax line-by-line works when your dictation engine handles identifier casing and operators deterministically.

How do you dictate an underscore or snake_case?

In Meosu, say "snake case max retries" to get max_retries. Camel case, kebab case, and screaming snake case work the same way. You can also spell symbols directly: "underscore", "open bracket", "double equals".