Ternary Transformer Weights
Ternary transformer weights refer to a quantization technique where model parameters are restricted to three discrete values (typically -1, 0, +1) to drastically reduce memory footprint and computational overhead while attempting to preserve reasoning capabilities. This approach is central to extreme quantization strategies for large language models (LLMs).
Key Developments
Ternary Bonsai 2 Evaluation
Recent evaluations have focused on Ternary Bonsai 2, a 27B-class reasoning model developed by prism-ml that utilizes ternary weights for extreme quantization.
- Model Specs: 27B parameters, supporting 1-bit and 2-bit quantization.
- Q1/Q2 Re-evaluation: Detailed benchmarking of Q1 and Q2 versions highlights performance trade-offs for local deployment. See Q2 Re-evaluation: Benchmarking Performance, Memory, Reasoning for granular data.
- Consumer Hardware Viability: Testing confirms feasibility on 16GB VRAM setups, making it a candidate for local LLM inference.
References
Bonsai-2-27B LLM Q1/Q2 Re-evaluation: Benchmarking Performance, Memory, Reasoning