Question 12
It requires both text and speaker embeddings as inputs.
The decoder input is compressed by a factor of 2 during training.
It transcribes speech into text like Whisper.
It can only generate speech in English.
It requires both text and speaker embeddings as inputs.
The decoder input is compressed by a factor of 2 during training.
It transcribes speech into text like Whisper.
It can only generate speech in English.
Correct answers
It requires both text and speaker embeddings as inputs.
The decoder input is compressed by a factor of 2 during training.
Question 12 of 19 in the IIT Madras BS Deep Learning Practice (Deep Learning Practice) Quiz 2 paper sat on 16 Mar 2025, in the January 2025 term (IIT M DEGREE AN EXAM QDB2 16 Mar 2025). It carries 3 marks.