Question 16
Why does DETR typically exhibit poor performance in detecting small objects compared to larger ones?
The CNN backbone used in DETR has a fixed receptive field that’s too large for small objects
The global self-attention mechanism in DETR tends to dilute the signal from small objects across all image locations
DETR’s loss function actively discards all detections smaller than 50x50 pixels
The object queries in DETR are programmed to ignore any object smaller than 10
DETR’s encoder-decoder architecture was specifically designed to only detect large objects