Dictation glossary

Speaker diarization

Speaker diarization means labeling a recording with who spoke when, usually as anonymous speaker tags rather than legal names. A diarization system finds turn boundaries and clusters voice characteristics so overlapping or sequential talkers get different labels. It sits beside speech-to-text: recognition writes the words, diarization attributes those words to a speaker track. Dictation of one person at a cursor rarely needs it; meeting transcripts often do.

In more detail

What is Speaker diarization?

What is speaker diarization in a pipeline? Audio is split into homogeneous segments, embeddings describe each segment's voice, and clustering assigns segment IDs. Names appear only if a later step matches those IDs to a roster. Two similar voices, a far-field mic, and crosstalk all raise the error rate. Diarization quality is measured with metrics such as diarization error rate, which is not word error rate. You can have perfect words with the wrong speaker, or the right speaker with mangled words.

Consumer 'meeting assistant' products advertise speaker labels because a wall of undivided text is hard to skim. That feature implies a recording of more than one person and usually a longer session than a dictation burst. It does not imply the app joined a video call; a laptop microphone in a conference room is enough to attempt diarization, and also enough to fail when everyone sits far away. If you only need to write as yourself, skip it.

For writers who speak

Why it matters for dictation

Writers confuse diarization with dictation and then feel cheated when a hotkey app does not tag colleagues. They are different jobs. If your work is notes from a room conversation, speaker turns help. If your work is drafting, speaker labels add complexity without helping the cursor. Ask for the feature only when you actually have multiple talkers on one recording.

In this product

How WhisperJot handles it

WhisperJot does not currently provide speaker labels or diarization. Meetings mode is microphone-only on-device transcription: a timestamped transcript of what your mic hears, with no bot, no call joining, and no AI summaries. Dictation remains a single-speaker writing workflow with local engines by default; Jot Cloud is opt-in for dictation only, never for meetings.

Questions

Straight answers.

What is speaker diarization?

Speaker diarization is the task of answering who spoke when on a recording. The system marks turn boundaries and assigns speaker IDs, often as Speaker 1 and Speaker 2 rather than real names. It is separate from speech-to-text, which produces the words. Meeting notes benefit; single-person dictation usually does not. Errors look like swapped labels or merged speakers, not like a misspelled word.

Is diarization the same as speaker recognition?

Not quite. Diarization clusters unknown talkers on one file. Speaker recognition, or identification, matches a voice to a known profile. You can diarize a meeting without knowing anyone's name, then optionally map IDs to a roster. Consumer apps blur the terms when they print first names. If identities matter, ask whether the product enrolls voices or only splits anonymous turns.

Do I need speaker diarization for voice dictation?

Almost never. Dictation assumes one author speaking into a focused app. There is no second speaker to label. Diarization starts to matter when you record a conversation and want the transcript to show turns. Buying a dictation hotkey and expecting meeting-style speaker columns mixes two products. Choose transcription features when the job is documenting a room, not when the job is writing.

One hotkey, any focused app.

Private local transcription by default, with an optional opt-in cloud engine.