Opening the paper…
Deep Learning for Computer Vision, Quiz 2
Why do very deep plain CNNs (without skip connections) sometimes show higher training error than shallower CNNs?
Why do very deep plain CNNs (without skip connections) sometimes show higher training error than shallower CNNs? MobileNetV1 reduces computation primarily by using: EfficientNet’s key scaling idea is: