Vision Transformer-Driven Multi-Modal Medical Imaging Framework for Early Diagnosis of Neurodegenerative Diseases

Authors

  • M. Padmapriya
  • Haritha Akkineni
  • M. Jeevana Sujitha
  • Lavanya Gottemukkala
  • Para Rajesh

DOI:

https://doi.org/10.51483/IJAIML.6.8s.2026.500-507

Keywords:

Vision Transformers, Multi-Modal Medical Imaging, Neurodegenerative Diseases, Alzheimer's Disease, Parkinson's Disease, Deep Learning, Medical Image Fusion.

Abstract

Accurate diagnosis of neurodegenerative diseases such as AD and PD in their early stages remains a critical challenge in modern medicine, due to the subtle and overlapping structural and functional changes in the brain during early stages. Conventional diagnostic methods, which rely on single modality imaging or subjective clinical evaluation, are not sensitive enough for early detection. We introduce a new multi-modal medical imaging framework based on Vision Transformers (ViTs) to improve early diagnosis of neurodegenerative diseases. The proposed framework is designed by integrating data from MRI and PET and exploiting the advantages of the strong self-attention of ViTs for modelling the overall contextual relationship and local structural anomalies. The multi-modal fusion strategy combines structural variations from MRI and metabolic activity from PET in a meaningful way, resulting in a comprehensive representation of brain pathology. The ViT-driven multi-modal framework outperforms the conventional Convolutional Neural Network (CNN)-based and single-modality methods in terms of classification accuracy, sensitivity and specificity, as demonstrated by experiments on publicly available datasets such as the Alzheimer's Disease Neuroimaging Initiative (ADNI) and the Parkinson's Progression Markers Initiative (PPMI). These results highlight the potential of Vision Transformers for multi-modal medical image analysis and provide a promising tool for early clinical intervention of neurodegenerative diseases.

Downloads

Published

2026-08-01

How to Cite

Padmapriya, M., Akkineni, H., Sujitha, M. J., Gottemukkala , L., & Rajesh, P. (2026). Vision Transformer-Driven Multi-Modal Medical Imaging Framework for Early Diagnosis of Neurodegenerative Diseases. International Journal of Artificial Intelligence and Machine Learning, 6(8s), 500–507. https://doi.org/10.51483/IJAIML.6.8s.2026.500-507