Question 1
In a speaker diarization pipeline that uses Whisper for transcription and an embedding model for speaker identification,you observe that one speaker's segments are consistently broken into multiple smaller segments, each assigned a different speaker label (e.g., SPEAKER 1, SPEAKER 3, SPEAKER 5). The transcribed text for this speaker is perfectly accurate. What is the MOST likely cause of this specific issue?
