Measure local dictation on your computer
Measure the time from stopping speech to inserted text, the time you spend correcting it, and memory pressure during a normal session. Keep the passage and settings consistent and repeat each condition. This page provides a method and blank worksheet, not published hardware benchmark results.
By WhisperJot · Documentation reviewed
Define the comparison
Choose a short everyday message and a longer passage containing terms you actually use. Keep a written reference for each. Record the exact engine and model, language, cleanup level, vocabulary, microphone, power mode, and whether the model was already loaded.
If you read aloud separately for each engine, delivery varies. Label this an end-to-end workflow comparison. A model-only accuracy comparison needs the identical audio file and matched processing settings. Do not compare those two methods as if they measured the same thing.
Time the whole usable result
- Run one warm-up dictation and record it separately from the timed attempts.
- For each attempt, measure from recording stop to the text appearing in the target document. If cleanup later replaces it, record final replacement time separately.
- Time your correction pass until the text matches the intended meaning and required spelling.
- Repeat at least three times per condition, alternating the order of engines. Keep all attempts, including failures.
- Record memory pressure and whether normal applications remain responsive. Report the median and range rather than only the fastest attempt.
Copy a results worksheet
Hardware / RAM: OS / WhisperJot version: Engine / exact model: Language / cleanup / vocabulary: Input / power mode / background apps: Passage / reference / trial count: Cold start or already loaded: Stop-to-insert times: Final cleanup times, if applicable: Correction times: Memory observations / failures: Method limitations:
| Condition | Insert time | Correction time | Memory / failures |
|---|---|---|---|
| Engine A, cold start | Not measured | Not measured | Not measured |
| Engine A, warm | Not measured | Not measured | Not measured |
| Engine B, warm | Not measured | Not measured | Not measured |
Interpret an 8 GB Mac result
WhisperJot's Mac installation guidance recommends 16 GB of RAM and describes Low-memory mode for 8 GB machines. That mode unloads models between sessions, so repeat the test with the mode recorded explicitly. A smaller idle footprint and repeated loading delays can coexist.
Choose the setup that produces usable text with acceptable waiting and editing time while your normal apps stay responsive. A faster raw transcript that takes much longer to correct may be a worse fit. Your findings apply to the machine and settings you recorded; they are not a universal engine ranking.
Questions
Is this a published WhisperJot benchmark?
No. It is a repeatable measurement procedure with an unfilled worksheet. No measured latency, accuracy, or memory result is asserted on this page.
Should I measure word error rate or editing time?
They answer different questions. Word error rate compares a transcript with a reference under an explicit normalization policy. Editing time measures practical effort and depends on the editor. Record both if useful and describe how each was obtained.
Sources and related guides
Product behavior is based on the documentation below, reviewed 2026-09-09. Examples are authored illustrations; measurement worksheets are for your own observations.
- WhisperJot Mac requirements and Low-memory mode
- Choosing a WhisperJot engine
- Word error rate explained