Independent of vault plugins
WhisperJot requires no Obsidian community plugin or transcription API key because it sends text to the application that currently has focus.
Obsidian has no built-in speech-to-text feature: its core Audio Recorder only attaches audio. You can use OS dictation, choose a community transcription plugin, or run WhisperJot and speak into the focused note without vault setup.
Obsidian's core Audio Recorder plugin creates an audio attachment but does not transcribe it. Community choices cover different setups: Whisper by nikdanilov sends audio to an OpenAI-Whisper-compatible API and needs an API key; Local Dictation and Speech Kit use offline local models without an API key; Voxtral Transcribe streams through the Mistral API and needs its key. OS dictation is another route. WhisperJot works outside the vault's plugin system and uses the same desktop shortcut in Obsidian and other apps.
WhisperJot requires no Obsidian community plugin or transcription API key because it sends text to the application that currently has focus.
Speech is handled on-device by default with local engines, and dictation remains available offline once a one-time model download is complete (about 500 MB on Mac, around 600 MB on Windows and Linux).
Move from an Obsidian note to a browser, editor, or email field and keep using the same system-level activation keys.
Add names, tags, and domain terms to custom vocabulary so recurring language from your knowledge system is easier to capture accurately.
Obsidian does not include native dictation. Its core Audio Recorder plugin can record and attach an audio file to a note, but it does not convert that recording to text. Live speech-to-text requires operating-system dictation, a community plugin with its own model or service setup, or a separate system-wide dictation application such as WhisperJot.
Whisper by nikdanilov uses an OpenAI-Whisper-compatible cloud API and requires an API key. Local Dictation and Speech Kit instead use offline models without an API key. Voxtral Transcribe provides streaming transcription through the Mistral API and therefore needs a Mistral key. Review each plugin's current documentation before selecting a workflow.
You can focus a note and use macOS Dictation, Windows voice typing, or WhisperJot without adding anything to the vault. WhisperJot runs as a desktop dictation tool, displays live partials in a floating HUD, and inserts the finished transcription at the active Obsidian cursor rather than attaching audio.
Offline options include community plugins such as Local Dictation and Speech Kit, which use local models without API keys. WhisperJot's local engines are also fully offline after a one-time model download (about 500 MB on Mac, around 600 MB on Windows and Linux) and process speech on-device by default. The Whisper and Voxtral Transcribe plugins listed here rely on cloud APIs instead.
With a local engine selected, transcription runs on your device by default, and WhisperJot keeps history in a local SQLite database. Its local models can work without a network connection after download. The optional Jot Cloud mode is opt-in behind explicit consent, operates only while selected, and never stores audio.
Sources, checked August 24, 2026: Whisper plugin for Obsidian, Local Dictation plugin. Vendor features change — check each product's documentation for current behavior.
Browse WhisperJot use cases, compare dictation apps, or check the system requirements.
Capture an idea by voice in Obsidian, then add a title, links and next action. Includes a sample note and a workflow for keeping dictation easy to review.
Read the workflow →Turn a checked voice-note transcript into a task list with owners, next steps and dates. Use an editable example and keep uncertain commitments visible.
Read the workflow →One $99/year plan includes unlimited dictation across your focused desktop apps.