Question 17
In a standard TTS pipeline involving an acoustic model and a vocoder, which component is primarily responsible for converting Mel-spectrograms into time-domain waveforms during the fine-tuning process?
The Vocoder.
The Duration Predictor.
The Phonemizer.
The Attention Mechanism.