Dictation glossary

Word error rate (WER)

Word error rate (WER) is a score of how many mistakes a speech recognizer made compared with a reference transcript. You count substitutions, insertions, and deletions, add them, and divide by the number of words in the reference. Lower is better; zero means the hypothesis matched the reference word for word. WER ignores punctuation in the usual formulation and can exceed 100 percent if the system inserts many extra words.

In more detail

What is WER?

What is word error rate in words? Align the machine hypothesis with a human reference. A substitution is a wrong word in the same slot. An insertion is an extra word the human never said. A deletion is a reference word the machine skipped. Word error rate equals substitutions plus insertions plus deletions, divided by the reference word count. If the reference is ten words and the recognizer substitutes one, inserts one, and deletes one, that is three errors over ten words, or a word error rate of 30 percent.

WER is a blunt tool. It treats every word equally, so missing 'not' costs the same as missing 'the,' which is a poor match for legal or medical risk. It also needs a reference, so you cannot compute it on live dictation unless you later type a gold transcript. Vendors quote WER on clean read speech that looks better than your open-plan office. Complement it with the errors you actually fix: names, jargon, and fillers. Custom vocabulary attacks a slice of WER that a bigger generic model may still miss.

For writers who speak

Why it matters for dictation

Accuracy marketing without a definition is noise. WER gives you a shared formula, and the worked example shows why a '5 percent' claim can hide painful name errors. For dictation, the practical test is still: how often do you reach for the keyboard? Use WER to compare engines on the same audio, not to crown a winner across mismatched test sets.

In this product

How WhisperJot handles it

WhisperJot does not publish a single word error rate, because WER depends on your microphone, accent, and domain. Jot Local Pro is the higher-accuracy on-device engine for long-form work; Jot Local is the faster short-utterance default. Custom vocabulary and text cleanup target the errors that show up in real dictation. Jot Cloud is opt-in if you want a hosted engine instead.

Questions

Straight answers.

What is word error rate?

Word error rate is a standard accuracy score for speech-to-text. You add the number of substituted, inserted, and deleted words, then divide by how many words are in the reference transcript. Lower percentages mean fewer alignment errors. The score can go above 100 percent when the system inserts heavily. It does not, by itself, tell you whether the remaining mistakes are harmless or catastrophic.

How do you calculate word error rate? Give an example.

Count substitutions, insertions, and deletions against a reference, add those three numbers, and divide by the reference length. Example: the reference has ten words; the hypothesis swaps one word, adds one word, and drops one word. Errors equal three, so word error rate is three divided by ten, or 30 percent. Punctuation is usually ignored. Always compare systems on the same reference audio.

Is a lower WER always a better dictation app?

Not by itself. Two engines with similar WER can feel different if one ruins names and the other ruins small function words. Latency, cleanup, and whether recognition is on-device also decide whether you will actually speak. Use WER on your own audio when you can, and treat vendor leaderboards as a starting point, not as a promise about your office.

One hotkey, any focused app.

Private local transcription by default, with an optional opt-in cloud engine.