What is speaker diarization in a pipeline? Audio is split into homogeneous segments, embeddings describe each segment's voice, and clustering assigns segment IDs. Names appear only if a later step matches those IDs to a roster. Two similar voices, a far-field mic, and crosstalk all raise the error rate. Diarization quality is measured with metrics such as diarization error rate, which is not word error rate. You can have perfect words with the wrong speaker, or the right speaker with mangled words.
Consumer 'meeting assistant' products advertise speaker labels because a wall of undivided text is hard to skim. That feature implies a recording of more than one person and usually a longer session than a dictation burst. It does not imply the app joined a video call; a laptop microphone in a conference room is enough to attempt diarization, and also enough to fail when everyone sits far away. If you only need to write as yourself, skip it.