Lexical Scoring Function

A mathematical function used in Information Retrieval (IR) to calculate the relevance of a document to a given query based on the frequency and distribution of terms. While modern search often relies on dense vector embeddings, lexical scoring functions remain critical for precision, interpretability, and handling rare or compound terms.

Core Concepts

  • Term Frequency (TF): Measures how often a term appears in a document.
  • Inverse Document Frequency (IDF): Measures how rare a term is across the entire corpus.
  • Sparse Representations: Relies on explicit term matching rather than semantic proximity.

BM25: The Standard Lexical Scorer

bm25 (Best Matching 25) is the most widely used lexical scoring function. It improves upon basic TF-IDF by:

  • Normalizing document length to prevent bias toward longer documents.
  • Using saturation functions to dampen the impact of extremely high term frequencies.

Despite the dominance of dense retrieval, lexical methods are seeing a resurgence in agentic-search contexts.

Comparison with Dense Retrieval

FeatureLexical (BM25)Dense (Vector)
BasisTerm frequency & raritySemantic embedding proximity
StrengthsExact matches, rare terms, speedSemantic understanding, synonymy
WeaknessesVocabulary mismatch, no semanticsComputationally heavier, less interpretable

References