Question 11
Regarding the SpeechT5 model for Text-to-Speech (TTS), which of the following statements are true? (Select ALL that apply)
It is an encoder-decoder Transformer model.
To generate speech for a specific voice, it requires a speaker embedding as an additional input.
HiFi-GAN is an integrated and mandatory part of the SpeechT5 model architecture.
It is pre-trained only on text data, similar to BERT.