Quiz Space

Deep Learning for Computer Vision · End Term · 31 Aug 2025 · May 2025 term · Set QIA3

Question 39: Consider the BLIP model architecture with an input (imag…

Question 39

+0.5 marksNumerical answer

Consider the BLIP model architecture with an input (image-text pair) as shown below:

For each of the entities given, enter the appropriate number from the figure that corresponds to it:
Based on the above data, answer the given subquestions.

Image-Text Contrastive Loss (LITC): ____________

Show answer

Correct answer: 10

Question 39 of 49 in the IIT Madras BS Deep Learning for Computer Vision (Deep Learning for Computer Vision) End Term paper sat on 31 Aug 2025, in the May 2025 term (IIT M IMPROVEMENT AN EXAM QIA3 31 Aug 2025). It carries 0.5 marks.

More questions from this paper

  1. Q1Why does DETR typically exhibit poor performance in detecting small objects compared to larger ones? Choose the best an…
  2. Q2Consider the following Pytorch code:\ x = torch.tensor([[-3.,2.,3.],[4.,-5.,7.],[20.,3.,2.]])\ print(torch.mean(x,dim=0…
  3. Q3Which one of the following statements is false?
  4. Q4What is the correct order of operations for processing an image through a Vision Transformer (ViT)?
  5. Q5Vector Quantized Variational Autoencoder (VQ-VAE) utilizes a discrete latent representation as opposed to continuous la…
  6. Q6Given two normalized embeddings from CLIP, I = [0.2, 0.5, 0.3, 0.4] for an image and T = [0.1, 0.6, 0.3, 0.7] for a tex…
  7. Q7Which one of the following statements regarding hyperparameter tuning is false?
  8. Q8Which one of the following statements is false? (Pick the most appropriate one.)
  9. Q9You are designing an edge detection pipeline for noisy grayscale images. For each of the following goals, choose the mo…
  10. Q10Figure question
  11. Q11Figure question
  12. Q12Figure question
  13. Q13In which one of the following applications would you use a many- to-one RNN architecture? Choose the most appropriate a…
  14. Q14Why might Segment Anything (SAM) be particularly useful in data annotation tasks compared to traditional segmentation m…
  15. Q15Which of the following statements are true? (Select all possible correct options)
  16. Q16Classifier-free guidance is a technique used in diffusion models to improve sample quality without the explicit use of …
  17. Q17Which of the following statements are true? (Select all possible correct options)
  18. Q18Which of the following statements about CLIP are TRUE? (Select ALL that apply)
  19. Q19Which one of the following statements is true:
  20. Q20Which of the following are examples of a high-pass filter?
  21. Q21Given is a 3 \times 3 8-bit grayscale image: \begin{bmatrix} 10 & 50 & 200 \ 70 & 100 & 150 \ 20 & 80 & 255 \end{bmatri…
  22. Q22Figure question
  23. Q23Figure question
  24. Q24Figure question
  25. Q25Figure question
  26. Q26What is the size of the feature map after applying two successive convolution operations with given parameters? Image s…
  27. Q27Based on the above data, answer the given subquestions.
  28. Q28Based on the above data, answer the given subquestions.
  29. Q29Element 1:_
  30. Q30Element 2:_
  31. Q31Element 3:_
  32. Q32Element 4:_
  33. Q33Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  34. Q34Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  35. Q35Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  36. Q36Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  37. Q37Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  38. Q38Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  39. Q40Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  40. Q41Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  41. Q42Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  42. Q43Consider the BLIP model architecture with an input (image-text pair) as shown below: For each of the entities given, en…
  43. Q44A four-dimensional input vector x = [4, -5, -7, 10] is passed to a hidden layer with a single neuron and an activation …
  44. Q45A four-dimensional input vector x = [4, -5, -7, 10] is passed to a hidden layer with a single neuron and an activation …
  45. Q46A four-dimensional input vector x = [4, -5, -7, 10] is passed to a hidden layer with a single neuron and an activation …
  46. Q47A four-dimensional input vector x = [4, -5, -7, 10] is passed to a hidden layer with a single neuron and an activation …
  47. Q48A four-dimensional input vector x = [4, -5, -7, 10] is passed to a hidden layer with a single neuron and an activation …
  48. Q49A four-dimensional input vector x = [4, -5, -7, 10] is passed to a hidden layer with a single neuron and an activation …