Multi-Cross Attention Network for Text, Image, And Audio Based Sentiment Analysis of Product Reviews
Keywords:
Multimodal sentiment analysis, cross-modal attention, MCAM, product review analysis, ALBERT, BiLSTM, DenseNet121, CBAM, MFCC, multimodal fusion.Abstract
Sentiment analysis tries to determine whether an opinion is positive, negative, or neutral. Most existing systems look only at text, yet a product review video usually carries three signals at once: the words that are spoken, the way they are spoken, and what is shown on screen. Using only one of these signals throws away clues that the other two could provide. This paper proposes a fusion module called the Multi-Cross Attention Mechanism (MCAM) that combines text, image, and audio features for sentiment classification of product review videos. Text features are learned with ALBERT and BiLSTM, image features with DenseNet121 and CBAM, and audio features with MFCC and BiLSTM. Rather than simply stacking these three feature sets together, MCAM allows each modality to exchange information with the other two so that the model can learn how they relate to one another. The combined representation is then passed through a fully connected layer and a Softmax layer that decides whether a review is positive, neutral, or negative. The framework is evaluated on a newly compiled dataset, MVSA++, containing close to ten thousand product review videos with matching text, image, and audio content. The complete model reaches 94.2% accuracy and a 92.7% F1-score, clearly ahead of models restricted to a single modality and ahead of a representative text-image baseline. These results indicate that combining text, image, and audio through cross-modal attention gives a fuller and more reliable picture of customer sentiment than relying on any one source alone.





