Question 26
Select all correct statements.
Soft attention can form a weighted combination of visual features.
Self-attention allows token representations to depend directly on other tokens.
Vision Transformers adapt transformer-style token processing to visual inputs.
Transformer-based segmentation is restricted to producing one class label for an entire image.