Internal Thoughts

Latent reasoning processes, intermediate activation states, or hidden cognitive layers within artificial neural networks (primarily llms) that precede final token generation. Unlike direct prompts or surface-level outputs, internal thoughts operate as unobservable or semi-observable mechanisms governing decision pathways, contextual synthesis, and value alignment before serialization into language.

Core Mechanisms

Emergent Internal Models

Recent mechanistic interpretability research has identified specific, discrete internal structures that emerge within large models, demonstrating that latent representations can encode functional algorithms distinct from the training data’s surface syntax.

References