Dictation & speech-to-text
glossary
This glossary defines speech-to-text and dictation terms in plain language so you can quote a one-paragraph answer, then read why the idea matters when you speak instead of type. Each definition is vendor-neutral; notes about WhisperJot sit in their own section on the term page.
Basics terms
Speech-to-text
Speech-to-text is the process of converting spoken audio into written words using a speech recognition model.
Read Speech-to-text →Voice dictation
Voice dictation is a way of writing in which you speak and software turns those words into editable text in an application.
Read Voice dictation →Voice typing
Voice typing means entering text by speaking instead of pressing keys, usually through a built-in control in an operating system or a single application.
Read Voice typing →Dictation vs transcription
Dictation vs transcription means the difference between speaking to produce editable text in a live writing workflow and converting recorded audio into a written record after the fact.
Read Dictation vs transcription →Push-to-talk
Push-to-talk means a recording or transmit channel that is active only while you hold a control, then stops when you release it.
Read Push-to-talk →Global hotkey
Global hotkey means a keyboard shortcut registered with the operating system so it works no matter which application is in front.
Read Global hotkey →Offline dictation
Offline dictation means speech-to-text that keeps working with no internet connection because the recognition model runs on your computer.
Read Offline dictation →Engines terms
Whisper
Whisper is an open-source family of automatic speech recognition models that convert recorded speech into text.
Read Whisper →Parakeet
Parakeet is a family of open speech recognition models from NVIDIA designed for fast automatic speech recognition.
Read Parakeet →On-device transcription
On-device transcription means speech recognition that runs on the computer or phone in front of you, using a model loaded in local memory.
Read On-device transcription →Cloud transcription
Cloud transcription means speech-to-text that runs on a remote server: your app uploads audio, a model in a data center produces text, and the words come back over the network.
Read Cloud transcription →Apple Neural Engine
Apple Neural Engine is a dedicated block of silicon on Apple chips that runs neural networks locally, alongside the CPU and GPU.
Read Neural Engine →Workflow terms
Text injection
Text injection means placing transcribed words into the application that currently has keyboard focus.
Read Text injection →Custom vocabulary
Custom vocabulary means a user-supplied list of words, names, and spellings that a speech-to-text system should prefer when it hears something similar.
Read Custom vocabulary →Filler-word removal
Filler-word removal means deleting spoken hedges — um, uh, like, you know — from a transcript so the remaining text reads more like writing than like a raw recording.
Read Filler-word removal →Live partial transcription
Live partial transcription means showing in-progress text while audio is still being captured, instead of waiting until you stop talking to reveal a finished transcript.
Read Live partials →Text cleanup
Text cleanup means editing a raw speech-to-text transcript so it looks like written language: punctuation, capitalization, filler-word removal, and user-defined replacements.
Read Text cleanup →Quality terms
Voice activity detection (VAD)
Voice activity detection (VAD) is a filter that decides whether a slice of audio contains speech or is silence, noise, or other sound.
Read VAD →Speaker diarization
Speaker diarization means labeling a recording with who spoke when, usually as anonymous speaker tags rather than legal names.
Read Speaker diarization →Word error rate (WER)
Word error rate (WER) is a score of how many mistakes a speech recognizer made compared with a reference transcript.
Read WER →Transcription latency
Transcription latency is the time between speech and usable text — first partial, or the finalized insert after you stop.
Read Latency →Speak instead of type.
WhisperJot types into the focused app and transcribes on your device by default.