Opening the paper…
Deep Learning for Computer Vision, Quiz 2
Why do very deep plain CNNs (without skip connections) sometimes show higher training error than shallower CNNs?
Why do very deep plain CNNs (without skip connections) sometimes show higher training error than shallower CNNs? In a standard ResNet-50 bottleneck block, the three convolutions are typically: In Inception/GoogLeNet modules, the main purpose of 1×1 convolutions is to: