
Deep Learning End Term: 13 September 2026, Set 1 (May 2026 term)
The IIT Madras BS Deep Learning (Deep Learning) End Term paper sat on 13 Sept 2026, in the May 2026 term, set 1: 19 questions for 50 marks in 180 minutes. Every question is below with its answer. Take it as a timed mock test to be marked, or read it through first.
- 19
- 50
- 180 min
- 13
- 2
- 4
Show answer
Correct answer: 79530
Question 2
The gradient at each layer depends on the gradient propagated from the layer above it.
A very large gradient can cause excessively large weight updates, potentially making training unstable.
Small gradients always cause the loss to decrease during training.
Show answer
Correct answers
The gradient at each layer depends on the gradient propagated from the layer above it.
A very large gradient can cause excessively large weight updates, potentially making training unstable.
Question 3
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 1
Question 4
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 4
Question 5
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 0
Question 6
Based on the above data, answer the given subquestions.
Yes
No
Show answer
Correct answer
Yes
Question 7
7
49
14
2048
Show answer
Correct answer
49
Question 8
Show answer
Correct answer: 2048
Question 9
Based on the above data, answer the given subquestions.
All three receive equal attention weight
Show answer
Correct answer
Question 10
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 0.544 (accepted within ±0.004)
Question 11
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 1.02 (accepted within ±0.02)
Question 12
Based on the above data, answer the given subquestions.
Let
Compute
and enter the sum of diagonal elements of the resulting matrix. Enter the answer correct to three decimal places.
Show answer
Correct answer: 0.0205 (accepted within ±0.0015)
Question 13
Based on the above data, answer the given subquestions.
Now let
Compute
and enter the sum of diagonal elements of the resulting matrix. Enter the answer correct to two decimal places.
Show answer
Correct answer: 15.175 (accepted within ±0.075)
Question 14
Based on the above data, answer the given subquestions.
Which statements correctly describes the two networks given in the previous two subquestions?
The first network exhibits exploding gradients while the second exhibits vanishing gradients.
The first network exhibits vanishing gradients while the second exhibits exploding gradients.
Show answer
Correct answer
The first network exhibits vanishing gradients while the second exhibits exploding gradients.
Question 15
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 3
Question 16
Based on the above data, answer the given subquestions.
Show answer
Correct answer: 0.268 (accepted within ±0.004)
Question 17
Show answer
Correct answer: 1801
Question 18
Which of the following statements are correct?
SVD requires an iterative gradient-based optimization procedure similar to the training of CBOW.
Show answer
Correct answers
Question 19
Show answer
Correct answer: 2.785 (accepted within ±0.003)