Thinking Tokens

Thinking Tokens refer to the intermediate computational steps or “chain-of-thought” sequences generated by Large Language Models (LLMs) during reasoning tasks. These tokens represent the model’s internal deliberation process, distinct from the final output. The efficiency of these tokens—measured by the ratio of reasoning effort to accuracy—is a critical metric for optimizing model performance and cost.

Key Developments & Evaluations

References