From Signal Processing to LLM-Integrated Speech Intelligence: A Narrative and Conceptual Review of Speech Processing
Keywords:
Speech Processing, Automatic Speech Recognition, Speech Synthesis, Deep Learning, Self-Supervised Learning, Multimodal Speech Processing, Speech Foundation ModelsAbstract
Audio and speech processing has evolved significantly over the past two decades, driven by advances in signal processing, statistical modeling, deep learning, and, more recently, large language models (LLMs). This paper presents a narrative and conceptual review of this evolution, highlighting major paradigm shifts that have shaped the field rather than attempting an exhaustive or systematic synthesis of the literature. We begin by examining classical signal processing approaches, which rely on handcrafted features and spectral analysis, followed by statistical methods that introduce probabilistic modeling for handling temporal variability and uncertainty. We then discuss the deep learning era, where end-to-end neural architectures enable hierarchical representation learning and substantially improve performance across a wide range of tasks, including speech recognition, synthesis, enhancement, and separation. Building on these developments, we analyze the emergence of self-supervised and foundation speech models, which leverage large-scale unlabeled data to learn general-purpose representations with improved robustness and cross-domain generalization. We further explore the integration of speech processing with large language models, emphasizing the transition from transcription-centric systems to multimodal speech intelligence. In this context, we highlight the fundamental challenge of bridging the modality gap between continuous acoustic signals and discrete textual representations. In addition to reviewing methodological advances, we identify several persistent challenges, including noise robustness, multi-speaker interference, data scarcity, evaluation limitations, and computational efficiency. We also discuss emerging research directions, such as unified speech foundation models, speech-native interaction with LLMs, and multimodal integration. This narrative review provides a unified conceptual perspective on the evolution of speech processing and outlines key directions for future research toward more integrated and intelligent speech systems.





