Leakage-Aware Multimodal Speech Emotion Recognition For Indian English: Sentence-Level Shortcuts Sarcasm and Uncertainty-Aware Fusion
Keywords:
Speech emotion recognition, Indian English, data leakage, sentence-disjoint evaluation, multimodal fusion, sarcasm detection, uncertainty-aware fusion, precision-weighted fusion, LoRA, cross-corpus evaluation.Abstract
Speech emotion recognition remains concentrated on Western-accented English corpora, leaving Indian English, spoken by hundreds of millions, largely unrepresented in resources for conversational agents, call-centre analytics, and affective monitoring systems. Standard speaker-disjoint splitting fails to prevent leakage when scripted corpora reuse a fixed sentence pool: identical transcripts recur across training and test sets, enabling silent lexical memorisation. No existing Indian-English resource includes a sarcastic class, leaving text–prosody conflict and uncertainty-aware fusion untested in this population. This study introduces the Real Sentiment Dataset (RSD), a 28-speaker, six-class Indian-English corpus with a sarcastic category, paired with a Lexical Shortcut Ratio diagnostic and leakage-controlled splits to quantify memorisation bias and benchmark fusion rigorously. Using LoRA-adapted RoBERTa and HuBERT encoders, we compare text-only, audio-only, concatenation, and precision-weighted fusion under speaker-disjoint, sentence-disjoint, and double-disjoint protocols, alongside cross-corpus transfer against IEMOCAP and a sarcasm modality-conflict analysis. Speaker-disjoint evaluation yields an inflated 100.00% text-only accuracy, collapsing to 40.83% once sentence overlap is removed; double-disjoint precision fusion performs best (62.47% WA), cross-corpus transfer remains near chance, and sarcasm raises branch disagreement without shifting fusion trust toward audio.These findings expose an underquantified evaluation flaw in scripted corpora and establish a reproducible foundation for Indian-English affective computing.





