Question 14
A 3.5-second audio file is sampled at 16000 Hz. It is fed into a 'Wav2Vec2Model'
("facebook/wav2vec2-base-960h"), which uses a CNN feature extractor that downsamples the input by a factor of 320. What will be the sequence length (i.e., the number of time steps) of the 'last hidden state' output?