Question 9
You transcribed an audio recording using Whisper and applied speaker diarization. However, you notice that the transcribed text is accurate, but the speaker labels frequently change mid- sentence, even when the same person is speaking.
Which of the following are definite causes of this issue?
The speaker segments are too short, causing the clustering model to misclassify speakers.
Whisper does not support speaker diarization, so the labels are randomly assigned.
The timestamps from Whisper do not align with the diarization output, leading to incorrect speaker switching.
The transcription model has misrecognized words, which affects speaker identification.