4-bit Quantization

4-bit quantization is a model compression technique that reduces the precision of neural network weights from standard 16-bit or 32-bit floating-point formats to 4-bit integers. This significantly decreases memory footprint and inference latency while attempting to preserve model accuracy.

Key Concepts

  • Precision Reduction: Maps continuous weight values to a discrete set of 16 possible values ().
  • Memory Efficiency: Reduces model size by ~75% compared to FP16, enabling deployment on consumer-grade hardware.
  • Performance Trade-off: May result in minor accuracy degradation, often mitigated by PTQ or QAT.

Recent Developments & Ecosystem

References

Source Notes