Cloud transcription is how most consumer voice features shipped for years. The client is a thin recorder; the intelligence lives behind an API. That split makes phones and old laptops feel fast, because they are not loading a 2 GB checkpoint. It also creates a queue: if the network is poor, dictation stutters. Providers differ on retention. Some say they discard audio after inference; some train on it. Those policies are contractual, not visible in the waveform, which is why consent and a clear default matter.
For dictation, cloud mode is a hardware escape hatch more than a different kind of speaking. You still press a hotkey and watch text appear. Accuracy depends on the hosted model, not on the word 'cloud.' Real-time factor can look excellent on a quiet fiber link and poor on hotel Wi-Fi. A product that is cloud-only cannot offer offline dictation. A product that is local-first can still include an optional cloud engine for machines under spec.