Ternary Transformer Weights

Ternary transformer weights refer to a quantization technique where model parameters are restricted to three discrete values (typically -1, 0, +1) to drastically reduce memory footprint and computational overhead while attempting to preserve reasoning capabilities. This approach is central to extreme quantization strategies for large language models (LLMs).

Key Developments

Ternary Bonsai 2 Evaluation

Recent evaluations have focused on Ternary Bonsai 2, a 27B-class reasoning model developed by prism-ml that utilizes ternary weights for extreme quantization.

References

Bonsai-2-27B LLM Q1/Q2 Re-evaluation: Benchmarking Performance, Memory, Reasoning