Sourced roundup · Updated 2026-09-05

Best speech recognition software for Linux
in 2026, by desktop workflow

We selected Linux speech recognition options by notes, simulated typing, hands-free control, and command-line transcription. Speech Note, nerd-dictation, Numen, and whisper.cpp cover different technical workflows; WhisperJot offers packaged desktop dictation in beta.

WhisperJot is our product. Competitor facts come from each vendor's public pages and were last verified on the dates shown.

Method

How we picked

Desktop input compatibility

Check X11, Wayland, and input permissions before choosing a typing tool. nerd-dictation offers several input backends; Numen uses dotool.

Model and setup requirements

Offline recognition still needs software and model setup. whisper.cpp accepts 16-bit WAV input; a CLI file workflow differs from dictating into the focused app.

Released tools versus upcoming features

WhisperJot Linux is beta. Canonical's Myna targets Ubuntu 26.10; do not treat an upcoming desktop feature as already shipped.

Options

Best speech recognition software for Linux options

Numbered for scanning, not as a ranking. Each pick is tagged with what it is best for.

  1. Speech Note

    Michal Kosciesza

    Best for: Offline notes and language tools

    Speech Note is a Linux notes application with offline speech to text, text to speech, and machine translation. Text and voice processing take place locally without a network connection. Its supported engines include Vosk, whisper.cpp, Faster Whisper, Coqui STT, and april-asr. The source repository uses the MPL-2.0 license, and the app is available through Flathub. Consider it when a dedicated notes workspace and multiple speech tools matter more than a minimal system-wide dictation shortcut.

    Open source; MPL-2.0

    Verified 2026-09-05

    Source

  2. nerd-dictation

    ideasman42

    Best for: Hackable offline typing on Linux

    nerd-dictation is a small, hackable Linux speech-to-text tool built on the Vosk API. Recognition works offline, and simulated typing can use xdotool, ydotool, dotool, or wtype. Those choices cover X11 and Wayland workflows, but you need to select and configure the input tool for your desktop. The project uses GPL-3.0. It suits people comfortable assembling their dictation setup and adjusting how recognized words reach the focused application, rather than expecting a packaged onboarding flow.

    Open source; GPL-3.0

    Verified 2026-09-05

    Source

  3. Numen

    Numen project

    Best for: Hands-free Linux control and typing

    Numen is voice control for hands-free Linux computing, using syllables and literal words for typing. Recognition runs locally with Vosk, and the setup downloads a Vosk binary and an English model of about 40 MB. Input is handled through dotool, which needs access to the input device interface and input group configuration. Numen supports X11, Wayland, and TTYs under an AGPLv3-only license. Consider it for a deliberate voice-control workflow that extends beyond ordinary prose dictation.

    Open source; AGPLv3 only

    Verified 2026-09-05

    Source

  4. whisper.cpp CLI

    ggml-org

    Best for: Local command-line file transcription

    whisper.cpp provides local speech recognition through a C and C++ implementation with CPU-only inference support. Its whisper-cli command runs on Linux and other platforms. This is a technical file-transcription choice rather than a ready-made desktop typing app. The documented CLI input is 16-bit WAV, with a conversion example using 16 kHz mono audio. Model downloads range from tiny at 75 MiB to large at 2.9 GiB. The project is released under the MIT license.

    Open source; MIT

    Verified 2026-09-05

    Source

  5. WhisperJot

    FastProducts LLC

    Best for: Packaged Linux dictation in beta

    WhisperJot's Linux app is in beta, with Ubuntu and Debian the most-tested path on modern 64-bit distributions. Local engines run on-device by default and work offline after the initial model download. Jot Local is around 600 MB on Linux; Jot Local Pro offers Whisper models for longer dictation. Custom vocabulary and voice commands are included in the same plan as stable Mac and Windows apps. Try the browser demo for a short dictation, while treating Linux desktop compatibility as a beta evaluation.

    $99/year, $30 every three months, or $199 lifetime; no free tier

    See WhisperJot pricing →
Summary

Best speech recognition software for Linux compared

Product Best for Price Runs where
Speech Note Offline notes and language tools Open source; MPL-2.0 Linux; offline notes application
nerd-dictation Hackable offline typing on Linux Open source; GPL-3.0 Linux; X11 and Wayland
Numen Hands-free Linux control and typing Open source; AGPLv3 only Linux; X11, Wayland, and TTYs
whisper.cpp CLI Local command-line file transcription Open source; MIT Linux and other platforms; command line
WhisperJot Packaged Linux dictation in beta $99/year, $30 every three months, or $199 lifetime; no free tier macOS 14+ Apple Silicon, Windows 10+ x64, Linux (beta)
Questions

Best speech recognition software for Linux FAQ

Can Linux speech recognition work offline?

Yes. Speech Note processes text and voice locally, nerd-dictation uses offline Vosk recognition, and Numen also recognizes speech locally. whisper.cpp provides local command-line transcription. WhisperJot's Linux beta works offline with its local engines after model setup. Choose the interface you need as well as the recognition engine.

Which Linux tools can type into other applications?

nerd-dictation simulates input using tools such as xdotool, ydotool, dotool, or wtype, depending on configuration. Numen uses dotool and supports X11, Wayland, and TTYs. WhisperJot offers a packaged dictation app in beta. Check desktop compatibility and input permissions; a file-transcription CLI is a different workflow from typing into the focused window.

Does Ubuntu have built-in dictation now?

Canonical's Myna project targets Ubuntu 26.10 and focuses on basic desktop dictation, with Wayland and GNOME as the primary validated environment. It uses local models and works offline after they are installed. As of September 5, 2026, do not treat that target as a shipped default Ubuntu feature; evaluate available tools separately.

Choose where your dictation runs.

Use local transcription by default, or explicitly opt in to Jot Cloud when you select it.