Quiz Space

Deep Learning for Computer Vision · End Term · 13 Apr 2025 · January 2025 term · Set 1-3

Question 40: Consider the BLIP model architecture with an input (imag…

Question 40

+0.54 marksNumerical answer

Consider the BLIP model architecture with an input (image-text pair) as shown below:
For each of the given entities, enter the appropriate number from the figure that corresponds to it:

Consider the BLIP model architecture with an input (image-text pair) as shown below:
For each of the given entities, enter the appropriate number from the figure that corresponds to it:
Based on the above data, answer the given subquestions.

Show answer

Correct answer: 2

Question 40 of 47 in the IIT Madras BS Deep Learning for Computer Vision (Deep Learning for Computer Vision) End Term paper sat on 13 Apr 2025, in the January 2025 term (IIT M IMPROVEMENT FN EXAM QIM2 13 Apr). It carries 0.54 marks.

More questions from this paper

  1. Q1Why does DETR typically exhibit poor performance in detecting small objects compared to larger ones? Choose the best an…
  2. Q2Which one of the following statements is false?
  3. Q3What is the correct order of operations for processing an image through a Vision Transformer (ViT)?
  4. Q4Vector Quantized Variational Autoencoder (VQ-VAE) utilizes a discrete latent representation as opposed to continuous la…
  5. Q5Given two normalized embeddings from CLIP, I = [0.2,−0.5, 0.3, 0.4] for an image and T = [−0.1, 0.6, −0.3,−0.7] for a t…
  6. Q6Which one of the following statements regarding hyperparameter tuning is false?
  7. Q7Which one of the following statements is false? (Pick the most appropriate one.)
  8. Q8Figure question
  9. Q9Figure question
  10. Q10Figure question
  11. Q11In which one of the following applications would you use a one- to-many RNN architecture? Choose the most appropriate a…
  12. Q12Why might Segment Anything (SAM) be particularly useful in data annotation tasks compared to traditional segmentation m…
  13. Q13Which of the following techniques help control the exploding or vanishing gradient problem in recurrent neural networks?
  14. Q14Which of the following statements are true? (Select all possible correct options)
  15. Q15Which of the following statements are true ? (Select all possible correct options)
  16. Q16Which of the following statements about CLIP are TRUE? (Select ALL that apply)
  17. Q17Which of the following statements are true?(Select ALL that apply)
  18. Q18Figure question
  19. Q19Figure question
  20. Q20Figure question
  21. Q21Consider a reverse process in a diffusion model where the goal is to reconstruct the original data from the noise. If t…
  22. Q22Figure question
  23. Q23Figure question
  24. Q24Figure question
  25. Q25Based on the above data, answer the given subquestions.
  26. Q26Based on the above data, answer the given subquestions.
  27. Q27Element 1:
  28. Q28Element 2:
  29. Q29Element 3:
  30. Q30Element 4:
  31. Q31Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  32. Q32Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  33. Q33Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  34. Q34Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  35. Q35Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  36. Q36Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  37. Q37Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  38. Q38Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  39. Q39Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  40. Q41Consider the BLIP model architecture with an input (image-text pair) as shown below:\ For each of the given entities, e…
  41. Q42Sigmoid
  42. Q43Linear
  43. Q44Indicator Function _
  44. Q45Softplus _
  45. Q46Relu _
  46. Q47Leaky-Relu