Whisper &
local speech models
Understand Whisper model sizes, runtime choices and Parakeet. Each explainer starts with a direct answer, keeps benchmark conditions attached to the numbers, and explains how WhisperJot uses the technology.
Models
Whisper model sizes
Compare Whisper model sizes using the official parameters, VRAM and relative speed table. Learn what English-only variants mean and how to choose.
parameters, VRAM and speed →Whisper large-v3
Learn what Whisper large-v3 changed from large-v2, how its language coverage and error-reduction claim are stated, and which license the project cites.
changes, languages and limits →Whisper large-v3-turbo
Understand Whisper large-v3-turbo: fewer decoder layers, reference speed, translation limits, compressed variants and its role in WhisperJot Jot Local Pro.
speed and tradeoffs →Distil-Whisper
Understand Distil-Whisper's English-only student models, decoder changes and reported speed, size and WER claims, including the distil-large-v3 checkpoint.
English speech distillation →Parakeet ASR
Compare Parakeet-TDT v2 and v3 languages, architecture, reported WER and throughput, timestamps and local runtimes, plus WhisperJot's English v2 engine.
TDT models and local runtimes →Runtimes
whisper.cpp
Explore whisper.cpp, its GGML model disk sizes, integer quantization and Apple acceleration options, plus how WhisperJot uses GGML on Windows and Linux.
GGML and local inference →faster-whisper
Read faster-whisper's CTranslate2 speed and memory benchmarks with their model and precision settings, Python requirements, and when this runtime fits.
CTranslate2 benchmarks →WhisperKit
Learn how WhisperKit runs Whisper through Core ML and the Apple Neural Engine, which Apple platforms it targets, and what compressed model names mean.
Whisper on Apple platforms →