ResNets
Residual Networks (ResNets) are a class of neural network architectures that utilize skip connections (or residual connections) to enable the training of extremely deep models. They address the degradation problem where accuracy saturates and then degrades as network depth increases, a phenomenon distinct from overfitting.
Core Concepts
- Residual Learning: Instead of learning a direct underlying mapping , the layers learn a residual function . The original mapping is recovered by .
- Skip Connections: Additive shortcuts that bypass one or more layers, allowing gradients to flow directly through the network during backpropagation.
- Solving Degradation: Mitigates the issue where deeper networks perform worse than shallower ones due to optimization difficulties.
- Shattered Gradients: Skip connections help maintain gradient magnitude and correlation across layers, preventing gradients from vanishing or exploding in very deep networks.
Key Insights from Welch Labs
- Historical Context: ResNets represent a pivotal moment in deep learning history, identified around 2015.
- The “Brilliant Hack”: The architecture is often described as an elegant solution to a fundamental optimization barrier in deep CNNs.
- Impact: It enabled the training of networks with hundreds or thousands of layers, significantly advancing state-of-the-art performance in computer vision.
Related Concepts
- Vanishing Gradient Problem
- Convolutional Neural Networks
- Backpropagation
- Identity Mapping