Quiz Space

Deep Learning Practice Quiz 2: 12 April 2026 (January 2026 term)

Question 1

+4 marksOne correct option

The tensor has shape: When using a pre-trained Transformer-based model (such as Wav2Vec 2.0) for a downstream classification task like Language Identification, what is the primary purpose of applying Mean Pooling across the sequence dimension?

The tensor  has shape:  When using a pre-trained Transformer-based model (such as Wav2Vec 2.0) for a downstream classifi
  1. A

    To reduce the sampling rate of the audio to make it compatible with standard deep learning layers.

  2. B

    To convert a variable-length sequence of hidden vectors into a single fixed- length representation that summarizes the entire utterance.

  3. C

    To ensure that the model focuses only on the most intense amplitude peaks of the speech signal.

  4. D

    To reverse the effects of the Convolutional Neural Network (CNN) encoder and return the data to the time domain.

Question 2

+4 marksOne correct option

When using Wav2Vec2Model to extract hidden states for a downstream task, you set in the model call. In PyTorch, what is the most efficient way to ensure the model does not calculate gradients or update weights during this feature extraction process, thereby saving memory and computation?

  1. A

    Use before passing the audio through the model.

  2. B

    Wrap the extraction code block with the : context manager.

  3. C

    Manually set each layer's attribute to False using a for loop before every forward pass.

  4. D

    Use the function on the output tensor to remove unnecessary dimensions.

Question 3

+4 marksOne correct option

When fine-tuning a Whisper model using the Hugging Face transformers library, which of the following statements best describes the role and behavior of the WhisperProcessor?

  1. A

    It is a standalone model that converts audio directly into text before passing it to the WhisperForConditionalGeneration class.

  2. B

    It is a wrapper class that combines a WhisperFeatureExtractor (for audio) and a WhisperTokenizer (for text) into a single object to simplify data preprocessing and padding.

  3. C

    It is primarily used to change the sampling_rate of the raw audio files to 16,000 Hz during the dataset.map() phase.

  4. D

    It is a specialized loss function that calculates the Word Error Rate (WER) by comparing the model's audio features directly against the ground-truth text label

15 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Deep Learning Practice Quiz 2 12 Apr 2026 paper

The IIT Madras BS Deep Learning Practice (Deep Learning Practice) Quiz 2 paper sat on 12 Apr 2026, in the January 2026 term: 18 questions for 52 marks in 120 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureDeep Learning Practice Quiz 2 12 Apr 2026 at a glance
TermJanuary 2026 term
SubjectDeep Learning Practice
Course codeBSDA5013
Questions18
Marks52
Duration120 min
MCQ10
Written5
MSQ3
Official paperDeep Learning Practice 06 Apr 26
Negative markingNo negative marking.
Updated

Same Quiz 2, other subjects

More Deep Learning Practice