Transformer-Based Deep Learning Model With Fusion-Based Reranking For Word-Level Sign Language Recognition

Authors

  • Maricel A. Esclamado

Keywords:

Transformer-based, Clustering, Fusion-based Reranking, Sign Language Recognition.

Abstract

Word-level sign language recognition remains challenging because of similarities among signs and the limited availability of balanced training data. This study proposes a Transformer-based deep learning model with fusion-based reranking for word-level sign language recognition using MediaPipe hand landmark sequences. A subset of the WLASL dataset was used in the analysis. Hand landmarks were extracted, preprocessed, and augmented before being used to train a Transformer encoder. To refine the predicted gloss rankings, clustering confidence scores obtained from K-means were combined with the Transformer's prediction probabilities through a fusion-based reranking strategy. Experimental results showed that reranking significantly improved Top-5, Top-10, and Top-20 accuracies, while Top-1 accuracy remained unchanged. These findings demonstrate that incorporating clustering information can effectively improve candidate ranking and enhance the overall performance of Transformer-based word-level sign language recognition.

Downloads

Published

2026-07-01

How to Cite

Esclamado, M. A. (2026). Transformer-Based Deep Learning Model With Fusion-Based Reranking For Word-Level Sign Language Recognition. International Journal of Artificial Intelligence and Machine Learning, 6(2), 21–27. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/786