uiz Space

May 2026 term · Deep Learning · BSCS3004

Deep Learning End Term: 13 September 2026, Set 1 (May 2026 term)

The IIT Madras BS Deep Learning (Deep Learning) End Term paper sat on 13 Sept 2026, in the May 2026 term, set 1: 19 questions for 50 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.

Questions
19
Marks
50
Duration
180 min
Numerical
13
MSQ
2
MCQ
4

Updated

Official paper: Deep Learning 13 Sep 26 (Session 2) · No negative marking.

Question 1

+4 marksNumerical answer
Show answer

Correct answer: 79530

Question 2

+3 marksOne or more correct options

Select all that apply.

  1. A

    The gradient at each layer depends on the gradient propagated from the layer above it.

  2. B
  3. C

    A very large gradient can cause excessively large weight updates, potentially making training unstable.

  4. D

    Small gradients always cause the loss to decrease during training.

Show answer

Correct answers

  • A

    The gradient at each layer depends on the gradient propagated from the layer above it.

  • B
  • C

    A very large gradient can cause excessively large weight updates, potentially making training unstable.

Question 3

+2 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 1

Question 4

+3 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 4

Question 5

+2 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 0

Question 6

+2 marksOne correct option

Based on the above data, answer the given subquestions.

  1. A

    Yes

  2. B

    No

Show answer

Correct answer

  • A

    Yes

Question 7

+2 marksOne correct option
  1. A

    7

  2. B

    49

  3. C

    14

  4. D

    2048

Show answer

Correct answer

  • B

    49

Question 8

+2 marksNumerical answer
Show answer

Correct answer: 2048

Question 9

+3 marksOne correct option

Based on the above data, answer the given subquestions.

  1. A
  2. B
  3. C
  4. D

    All three receive equal attention weight

Show answer

Correct answer

  • A

Question 10

+2 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 0.544 (accepted within ±0.004)

Question 11

+3 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 1.02 (accepted within ±0.02)

Question 12

+3 marksNumerical answer

Based on the above data, answer the given subquestions.

Let

Compute

and enter the sum of diagonal elements of the resulting matrix. Enter the answer correct to three decimal places.

Show answer

Correct answer: 0.0205 (accepted within ±0.0015)

Question 13

+3 marksNumerical answer

Based on the above data, answer the given subquestions.

Now let

Compute

and enter the sum of diagonal elements of the resulting matrix. Enter the answer correct to two decimal places.

Show answer

Correct answer: 15.175 (accepted within ±0.075)

Question 14

+2 marksOne correct option

Based on the above data, answer the given subquestions.

Which statements correctly describes the two networks given in the previous two subquestions?

  1. A
  2. B

    The first network exhibits exploding gradients while the second exhibits vanishing gradients.

  3. C

    The first network exhibits vanishing gradients while the second exhibits exploding gradients.

  4. D
Show answer

Correct answer

  • C

    The first network exhibits vanishing gradients while the second exhibits exploding gradients.

Question 15

+3 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 3

Question 16

+3 marksNumerical answer

Based on the above data, answer the given subquestions.

Show answer

Correct answer: 0.268 (accepted within ±0.004)

Question 17

+2 marksNumerical answer
Show answer

Correct answer: 1801

Question 18

+3 marksOne or more correct options

Which of the following statements are correct?

Select all that apply.

  1. A
  2. B
  3. C

    SVD requires an iterative gradient-based optimization procedure similar to the training of CBOW.

  4. D
Show answer

Correct answers

  • A
  • B

Question 19

+3 marksNumerical answer
Show answer

Correct answer: 2.785 (accepted within ±0.003)