Shattered Gradients
Shattered Gradients refers to the phenomenon where gradient magnitudes decay exponentially as they propagate backward through deep neural networks, particularly in early architectures like deep Convolutional Neural Network prior to the widespread adoption of residual connections. This effect hinders training stability and convergence in very deep models.
Key Characteristics
- Exponential Decay: Gradient norms shrink rapidly with depth, causing early layers to receive negligible update signals.
- Vanishing Gradient Relation: Often conflated with or exacerbated by the Vanishing Gradient Problem, but specifically highlights the fragmentation of gradient information across layers.
- Impact: Leads to poor learning in deep networks, necessitating architectural innovations to maintain gradient flow.
Solutions and Mitigations
- Residual Connections (Skip Connections): The primary solution introduced by ResNets, which allow gradients to flow directly through the network, bypassing non-linear transformations.
- Batch Normalization: Helps stabilize layer inputs, reducing internal covariate shift and aiding gradient propagation.
- Careful Initialization: Techniques like Xavier Initialization or He Initialization help maintain variance in activations and gradients.