A Hybrid Quantum Re-Uploading Framework for Multimodal Emotion Recognition
Keywords:
Quantum Machine Learning; Multimodal Emotion Recognition; Quantum Re-Uploading; Affective Computing; Hybrid Quantum-Classical Models.Abstract
Multimodal emotion recognition can benefit from combining complementary information from facial images, speech, and text. However, integrating heterogeneous modalities obtained from independently collected datasets introduces challenges in data alignment and label consistency. This study presents a hybrid quantum–classical framework that combines modality-specific classical encoders with a quantum re-uploading variational classifier for six-class emotion recognition. Facial images, speech samples, and textual samples were obtained from separate publicly available datasets and aligned at the emotion-label level to construct a unified multimodal experimental setting. Image, speech, and text encoders generated modality-specific representations that were fused through a classical neural network and projected into a six-dimensional quantum input space. A six-qubit quantum circuit with three re-uploading layers repeatedly encoded the fused features before six-class classification. On a test set of 6,347 samples, the classical Image–Speech–Text fusion model achieved 73.94% accuracy and a 70.64% macro F1-score, whereas the proposed quantum re-uploading model achieved 72.57% accuracy and a 68.76% macro F1-score. The quantum model used 8,706 trainable parameters compared with 15,440,006 for the classical multimodal model, representing a 99.94% reduction in trainable parameters. The results indicate that quantum re-uploading can provide competitive performance for multimodal emotion recognition while highlighting the limitations of label-level alignment across independently collected datasets.





