Convolutional Neural Networks
Convolutional Neural Networks (CNNs) are a class of deep learning models primarily designed for processing grid-like topology data, such as images. They utilize convolutional layers to automatically and adaptively learn spatial hierarchies of features.
Core Components
- Convolutional Layers: Apply learnable filters (kernels) to input data to produce feature maps.
- Activation Functions: Non-linearities (e.g., ReLU) applied after convolution to introduce non-linearity.
- Pooling Layers: Downsample feature maps to reduce dimensionality and computational load (e.g., Max Pooling).
- Fully Connected Layers: Typically used at the end for classification or regression tasks.
Evolution and Key Architectures
Early CNNs
- LeNet-5: Pioneered the use of convolutional layers for digit recognition.
- AlexNet: Demonstrated the power of deep CNNs using GPUs and ReLU activations.
- VGGNet: Explored the impact of network depth using small (3x3) convolution filters.
ResNets: Solving Deep CNN Degradation and Shattered Gradients with Skip Connections
As networks became deeper, performance saturated and then degraded rapidly. This was not due to overfitting but to the degradation problem, where deeper networks become harder to optimize.
- The Degradation Problem: Increasing depth leads to higher training error, contrary to expectations.
- Shattered Gradients: In very deep networks, gradients can vanish or explode, making learning unstable.
- Residual Learning: Introduced by ResNets: Solving Deep CNN Degradation and Shattered Gradients with Skip Connections.
- Skip Connections (Residual Blocks): Allow gradients to flow directly through the network, bypassing non-linear transformations. This enables the training of extremely deep networks (e.g., ResNet-152).
- Identity Mapping: The residual function is easier to learn than the direct mapping .
Modern Developments
- Inception Modules: Use parallel convolutions of different sizes to capture multi-scale features.
- Dense Connections: Each layer receives inputs from all preceding layers (DenseNet).
- Attention Mechanisms: Integrated into CNNs (e.g., Vision Transformers) to weigh the importance of different parts of the input.