Quiz Space

Deep Learning Practice · End Term · 13 Sept 2026 · May 2026 term · Set 1

Question 30: SRGAN performs photo-realistic super-resolution using a …

Question 30

+2 marksOne or more correct options

SRGAN performs photo-realistic super-resolution using a GAN. Which of the following statements are correct?

Select all that apply.

  1. A

    Its perceptual loss combines a content loss and an adversarial loss, rather than relying on pixel-wise MSE alone.

  2. B

    The content loss is computed on feature maps of a pre-trained VGG network (perceptual similarity), not purely in pixel space.

  3. C

    Purely MSE-optimized solutions tend to be overly smooth because they approximate a pixel-wise average of many plausible HR solutions, losing high-frequency texture.

  4. D

    The adversarial term makes SRGAN maximize PSNR, which is why it always achieves a higher PSNR than the MSE-based SRResNet.

Show answer

Correct answers

  • A

    Its perceptual loss combines a content loss and an adversarial loss, rather than relying on pixel-wise MSE alone.

  • B

    The content loss is computed on feature maps of a pre-trained VGG network (perceptual similarity), not purely in pixel space.

  • C

    Purely MSE-optimized solutions tend to be overly smooth because they approximate a pixel-wise average of many plausible HR solutions, losing high-frequency texture.

Question 30 of 31 in the IIT Madras BS Deep Learning Practice (Deep Learning Practice) End Term paper sat on 13 Sept 2026, in the May 2026 term (Deep Learning Practice 13 Sep 26 (Session 2)). It carries 2 marks.

More questions from this paper

  1. Q1Based on the above data, answer the given subquestions.
  2. Q2Based on the above data, answer the given subquestions.
  3. Q3Using the actual flattened feature size from previous question number 2, how many learnable parameters (weights + bias)…
  4. Q4Which of the following statements about the (corrected) model are true?
  5. Q5Figure question
  6. Q6Figure question
  7. Q7Which layer is responsible for the decoder failing to reach the target 120×120 output, and why?
  8. Q8Figure question
  9. Q9Figure question
  10. Q10Figure question
  11. Q11What is the parameter compression ratio (standard ÷ depthwise-separable), rounded to 2 decimal places?
  12. Q12Figure question
  13. Q13Which of the following statements correctly identify a real bug in the given code ?
  14. Q14Based on the above data, answer the given subquestions.
  15. Q15Based on the above data, answer the given subquestions.
  16. Q16A convolutional layer receives a feature map with 16 input channels and applies 32 filters, each of spatial size 5 × 5.…
  17. Q17A convolutional layer is applied to a 32 × 32 input using a 5 × 5 filter with zero-padding of 2 and stride 2. Using the…
  18. Q18The final classification head applies a softmax over the pre-activation scores (logits). For a 3-class problem, the log…
  19. Q19Following the VGG design philosophy of replacing a single large-kernel convolution with a cascade of small 3 × 3, strid…
  20. Q20Figure question
  21. Q21The decoder of a depth-estimation network upsamples feature maps using a transposed convolution ("up-convolution") with…
  22. Q22In the original U-Net used for dense prediction (and adaptable to depth estimation), each encoder stage applies two suc…
  23. Q23The SRGAN discriminator is a stack of 8 convolutional layers, all using 3 × 3 kernels with padding 1, and strides alter…
  24. Q24Efficient restoration networks (e.g., LaKDNet) replace standard convolutions with depth-wise separable convolutions (a …
  25. Q25Compared with the sigmoid activation, which of the following are correct reasons for preferring the ReLU activation in …
  26. Q26VGGNet replaces large convolution kernels with stacks of small 3 × 3 convolutions. Which of the following statements co…
  27. Q27Which of the following statements correctly describe the Region Proposal Network (RPN) in Faster R-CNN?
  28. Q28Consider the multi-task loss used to train the original YOLOv1 model. Which of the following statements about it are co…
  29. Q29Regarding the channel-attention module in CBAM, which of the following are correct?
  30. Q31In the multi-scale deep network for single-image depth estimation (Eigen et al.), the architecture is split into a coar…