Temperature Parameter
The Temperature Parameter controls the randomness of predictions by a large-language-model (LLM). It scales the logits before the softmax function, influencing the probability distribution of the next token.
Key Mechanics
- Low Temperature (< 1.0): Makes the model more deterministic and confident. Reduces creativity and increases the likelihood of repeating common patterns.
- High Temperature (> 1.0): Increases randomness and creativity. Can lead to more diverse outputs but may result in incoherent or nonsensical text.
- Temperature = 0: Deterministic sampling (greedy decoding). The model always picks the most likely next token.
Impact on Quantized Models
Quantization affects how temperature influences output stability. Lower-bit quantizations may exhibit higher sensitivity to temperature changes due to reduced precision in weight representations.
Recent Benchmarks: Bonsai-2-27B
Recent evaluations of the Ternary-Bonsai-2-27B-gguf model highlight the interaction between quantization levels and temperature settings:
- Q2 Re-evaluation: Benchmarking Performance, Memory, Reasoning
- Re-evaluation of Q1 vs Q2 quantized versions from Prism ML.
- Focus on performance, memory usage, and reasoning capabilities in a 16GB local LLM setup.
- Analysis by Luke’s Dev Lab indicates that optimal temperature settings vary significantly between Q1 and Q2 quantizations due to differences in weight distribution and precision loss.
Best Practices
- Use lower temperatures for factual, code, or logical tasks.
- Use higher temperatures for creative writing, brainstorming, or open-ended dialogue.
- Always test temperature sensitivity when deploying new quantized models to ensure consistent behavior.
References
Bonsai-2-27B LLM Q1/Q2 Re-evaluation: Benchmarking Performance, Memory, Reasoning