Question 1
Why does DETR typically exhibit poor performance in detecting small objects compared to larger ones? Choose the best answer in your opinion.
The CNN backbone used in DETR has a fixed receptive field that’s too large forsmall objects.
The global self-attention mechanism in DETR tends to dilute the signal fromsmall objects across all image locations.
DETR’s loss function actively discards all detections smaller than 50 × 50 pixels.
The object queries in DETR are programmed to ignore any object smaller than10% of the image size.
DETR’s encoder-decoder architecture was specifically designed to only detectobjects larger than a cat.