MGT-AATA-Net: Multi-Granular Token and Attention Mechanism Based Facial Expression Recognition
Keywords:
Multi Granular Tokenisation Alignment-Aware Temporal Attention (MGT-AATA-Net), Facial Expression Recognition (FER), Bidirectional Long Short-Term Memory (BiLSTM), Alignment-Aware Temporal Attention (AATA), Residual Network-18 (ResNet-18), Occlusion-Adaptive Dynamic Routing (OADR).Abstract
Facial Expression Recognition (FER) remains a challenging problem in affective computing due to highly heterogeneous spatial characteristics of facial expressions, where emotion-discriminative information may range from global facial configurations to localized muscle movements and subtle micro-textural variations. Conventional convolutional and attention-based FER models predominantly focus on single-scale or visually salient representations, which can overlook low-contrast yet semantically critical facial cues and consequently limit recognition of difficult emotion categories. To address these limitations, this paper proposes Multi Granular Tokenization Alignment-Aware Temporal Attention Network (MGT-AATA-Net), a novel deep learning framework that integrates multi-granular representation learning with alignment-aware attention for robust and discriminative FER. The proposed architecture decomposes facial representations into macro, meso, and micro granular tokens, enabling simultaneous modelling of global facial geometry, intermediate spatial relationships, and fine-grained expression-specific textures. This hierarchical representation is designed to preserve complementary information across spatial frequency levels and thereby improve sensitivity to subtle and localized affective cues. In addition, an Alignment-Aware Temporal Attention (AATA) mechanism incorporates semantic alignment priors derived from the Weighted Directional Centroid Ranking Function (WDCRF) into the attention logits. Unlike conventional self-attention, which may preferentially allocate attention to visually dominant but emotionally irrelevant regions, the proposed mechanism explicitly biases feature interactions toward semantically informative facial regions while suppressing spurious salient responses. On the other hand Occlusion-Adaptive Dynamic Routing (OADR) was also applied to address the occlusion condition during FER process.
The proposed framework is evaluated with RAF-DB dataset, experimental results demonstrate that MGT-AATA-Net achieves overall accuracy of 93.0% and macro-F1 of 91.6%. The robustness of the proposed model was also evaluated through occlusion study, ablation study, GRAD-CAM visibility and ROC-AUC performance study and found satisfactory. These findings established multi-granular semantic alignment as a promising direction for developing more robust FER system capable of capturing both coarse facial geometry and subtle expression-level cues.





