Convolutional Neural Networks

Convolutional Neural Networks (CNNs) are a class of deep learning models primarily designed for processing grid-like topology data, such as images. They utilize convolutional layers to automatically and adaptively learn spatial hierarchies of features.

Core Components

  • Convolutional Layers: Apply learnable filters (kernels) to input data to produce feature maps.
  • Activation Functions: Non-linearities (e.g., ReLU) applied after convolution to introduce non-linearity.
  • Pooling Layers: Downsample feature maps to reduce dimensionality and computational load (e.g., Max Pooling).
  • Fully Connected Layers: Typically used at the end for classification or regression tasks.

Evolution and Key Architectures

Early CNNs

  • LeNet-5: Pioneered the use of convolutional layers for digit recognition.
  • AlexNet: Demonstrated the power of deep CNNs using GPUs and ReLU activations.
  • VGGNet: Explored the impact of network depth using small (3x3) convolution filters.

ResNets: Solving Deep CNN Degradation and Shattered Gradients with Skip Connections

As networks became deeper, performance saturated and then degraded rapidly. This was not due to overfitting but to the degradation problem, where deeper networks become harder to optimize.

Modern Developments

  • Inception Modules: Use parallel convolutions of different sizes to capture multi-scale features.
  • Dense Connections: Each layer receives inputs from all preceding layers (DenseNet).
  • Attention Mechanisms: Integrated into CNNs (e.g., Vision Transformers) to weigh the importance of different parts of the input.

References