Desktop input compatibility
Check X11, Wayland, and input permissions before choosing a typing tool. nerd-dictation offers several input backends; Numen uses dotool.
We selected Linux speech recognition options by notes, simulated typing, hands-free control, and command-line transcription. Speech Note, nerd-dictation, Numen, and whisper.cpp cover different technical workflows; WhisperJot offers packaged desktop dictation in beta.
WhisperJot is our product. Competitor facts come from each vendor's public pages and were last verified on the dates shown.
Check X11, Wayland, and input permissions before choosing a typing tool. nerd-dictation offers several input backends; Numen uses dotool.
Offline recognition still needs software and model setup. whisper.cpp accepts 16-bit WAV input; a CLI file workflow differs from dictating into the focused app.
WhisperJot Linux is beta. Canonical's Myna targets Ubuntu 26.10; do not treat an upcoming desktop feature as already shipped.
Numbered for scanning, not as a ranking. Each pick is tagged with what it is best for.
Michal Kosciesza
Best for: Offline notes and language tools
Speech Note is a Linux notes application with offline speech to text, text to speech, and machine translation. Text and voice processing take place locally without a network connection. Its supported engines include Vosk, whisper.cpp, Faster Whisper, Coqui STT, and april-asr. The source repository uses the MPL-2.0 license, and the app is available through Flathub. Consider it when a dedicated notes workspace and multiple speech tools matter more than a minimal system-wide dictation shortcut.
Open source; MPL-2.0
Verified 2026-09-05
ideasman42
Best for: Hackable offline typing on Linux
nerd-dictation is a small, hackable Linux speech-to-text tool built on the Vosk API. Recognition works offline, and simulated typing can use xdotool, ydotool, dotool, or wtype. Those choices cover X11 and Wayland workflows, but you need to select and configure the input tool for your desktop. The project uses GPL-3.0. It suits people comfortable assembling their dictation setup and adjusting how recognized words reach the focused application, rather than expecting a packaged onboarding flow.
Open source; GPL-3.0
Verified 2026-09-05
Numen project
Best for: Hands-free Linux control and typing
Numen is voice control for hands-free Linux computing, using syllables and literal words for typing. Recognition runs locally with Vosk, and the setup downloads a Vosk binary and an English model of about 40 MB. Input is handled through dotool, which needs access to the input device interface and input group configuration. Numen supports X11, Wayland, and TTYs under an AGPLv3-only license. Consider it for a deliberate voice-control workflow that extends beyond ordinary prose dictation.
Open source; AGPLv3 only
Verified 2026-09-05
ggml-org
Best for: Local command-line file transcription
whisper.cpp provides local speech recognition through a C and C++ implementation with CPU-only inference support. Its whisper-cli command runs on Linux and other platforms. This is a technical file-transcription choice rather than a ready-made desktop typing app. The documented CLI input is 16-bit WAV, with a conversion example using 16 kHz mono audio. Model downloads range from tiny at 75 MiB to large at 2.9 GiB. The project is released under the MIT license.
Open source; MIT
Verified 2026-09-05
FastProducts LLC
Best for: Packaged Linux dictation in beta
WhisperJot's Linux app is in beta, with Ubuntu and Debian the most-tested path on modern 64-bit distributions. Local engines run on-device by default and work offline after the initial model download. Jot Local is around 600 MB on Linux; Jot Local Pro offers Whisper models for longer dictation. Custom vocabulary and voice commands are included in the same plan as stable Mac and Windows apps. Try the browser demo for a short dictation, while treating Linux desktop compatibility as a beta evaluation.
$99/year, $30 every three months, or $199 lifetime; no free tier
See WhisperJot pricing →| Product | Best for | Price | Runs where |
|---|---|---|---|
| Speech Note | Offline notes and language tools | Open source; MPL-2.0 | Linux; offline notes application |
| nerd-dictation | Hackable offline typing on Linux | Open source; GPL-3.0 | Linux; X11 and Wayland |
| Numen | Hands-free Linux control and typing | Open source; AGPLv3 only | Linux; X11, Wayland, and TTYs |
| whisper.cpp CLI | Local command-line file transcription | Open source; MIT | Linux and other platforms; command line |
| WhisperJot | Packaged Linux dictation in beta | $99/year, $30 every three months, or $199 lifetime; no free tier | macOS 14+ Apple Silicon, Windows 10+ x64, Linux (beta) |
Yes. Speech Note processes text and voice locally, nerd-dictation uses offline Vosk recognition, and Numen also recognizes speech locally. whisper.cpp provides local command-line transcription. WhisperJot's Linux beta works offline with its local engines after model setup. Choose the interface you need as well as the recognition engine.
nerd-dictation simulates input using tools such as xdotool, ydotool, dotool, or wtype, depending on configuration. Numen uses dotool and supports X11, Wayland, and TTYs. WhisperJot offers a packaged dictation app in beta. Check desktop compatibility and input permissions; a file-transcription CLI is a different workflow from typing into the focused window.
Canonical's Myna project targets Ubuntu 26.10 and focuses on basic desktop dictation, with Wayland and GNOME as the primary validated environment. It uses local models and works offline after they are installed. As of September 5, 2026, do not treat that target as a shipped default Ubuntu feature; evaluate available tools separately.
See sourced one-on-one pages in alternatives or the broader dictation app comparison.
Use local transcription by default, or explicitly opt in to Jot Cloud when you select it.