Token-by-Text Generation

Token-by-token text generation refers to the autoregressive process where a model predicts the next token in a sequence based on previous tokens, iteratively constructing the final output. This method is the standard inference mechanism for large-language-models (LLMs).

Core Mechanism

  • Autoregressive Prediction: The model computes a probability distribution over the vocabulary for the next token.
  • Sampling/Decoding: A token is selected (via greedy, sampling, or beam search) and appended to the context.
  • Latency Implications: Sequential dependency creates inherent latency, as each token must be generated before the next can be computed.

Optimization Context: On-Device Efficiency

Recent developments challenge the necessity of massive LLMs for specific tasks like Function Calling.

References