Quiz Space

Deep Learning for Computer Vision End Term: 22 December 2024 (September 2024 term)

Question 1

+2 marksOne correct option

What is the correct order of operations for processing an image through a Vision Transformer (ViT)?

  1. A

    Image patching → Positional embedding → Linear projection of flattened patches → Transformer encoder→ Classification head

  2. B

    Image patching → Linear projection of flattened patches → Positional embedding → Transformer encoder→ Classification head

  3. C

    Positional embedding → Image patching → Linear projection of flattened patches → Transformer encoder→ Classification head

  4. D

    Linear projection of Images → Image patching → Positional embedding → Transformer encoder → Classification head

  5. E

    Linear projection of Images → Image patching → Positional embedding → Transformer encoder→ Transformer decoder→ Classification head

Also asked in End Term 31 Aug 2025

Question 2

+2 marksOne correct option

Vector Quantized Variational Autoencoder (VQ-VAE) utilizes a discrete latent representation as opposed to continuous latent spaces used in traditional VAEs. One of the key components of a VQ- VAE is the codebook, which consists of a set of learnable vectors. What is the primary role of the codebook in VQ-VAE?

  1. A

    It performs the non-linear transformation of the input data to a higher dimensional space.

  2. B

    It regularizes the encoder by penalizing complex encodings.

  3. C

    It provides a finite set of vectors which the encoder’s outputs are mapped to, effectively quantizing the latent space.

  4. D

    It decodes the quantized vectors back into the reconstructed input space.

Also asked in End Term 31 Aug 2025

Question 3

+2 marksOne correct option

Consider a reverse process in a diffusion model where the goal is to reconstruct the original data from the noise. If the model correctly reduces the variance of the noise by 0.02 in each reverse step, and starts with a noise variance of 1.0 at timestep T = 50, how many steps are required to reduce the noise variance to 0.1?

  1. A

    45

  2. B

    50

  3. C

    40

  4. D

    30

22 more questions in this paper

Sign in with Google — it is free — to see every question with its answer and explanation, practise it in learning mode, or take it as a timed mock test.

More on the Deep Learning for Computer Vision End Term 22 Dec 2024 paper

The IIT Madras BS Deep Learning for Computer Vision (Deep Learning for Computer Vision) End Term paper sat on 22 Dec 2024, in the September 2024 term: 25 questions for 50 marks in 180 minutes. The first 3 questions are below. Sign in with Google — it is free — to see the whole paper with its answers and explanations, in learning mode or as a timed mock test.

FeatureDeep Learning for Computer Vision End Term 22 Dec 2024 at a glance
TermSeptember 2024 term
SubjectDeep Learning for Computer Vision
Course codeBSDA5006
Questions25
Marks50
Duration180 min
MCQ15
MSQ2
Numerical8
Official paperIIT M DEGREE AN EXAM QDB4 22 Dec 2024
Negative markingNo negative marking.
Updated

Same End Term, other subjects

More Deep Learning for Computer Vision