Unified Audio-Text Model

Unified Audio-Text Models are multimodal architectures capable of processing, understanding, and generating both audio and text data within a single framework. Unlike traditional pipelines that separate speech recognition (ASR) from text generation, these models natively handle audio tokens alongside text tokens, enabling seamless interaction between spoken and written modalities.

Key Characteristics

Notable Implementations

NVIDIA Audex-2B

A recent addition to the nemotron family, Audex-2B represents a shift toward compact, efficient unified models.

References