Cross-Modal Intelligence Framework for Vision-Language Integration in Autonomous Driving Applications

Authors

  • M. Anitha
  • Dr.K. Sreedevi
  • Dr.K.R. Sowmya
  • Dr.R. Udayakumar
  • Dr Pennada Siva Satya Prasad

Keywords:

Cross-Modal Learning, Vision-Language Models, Autonomous Driving, Contrastive Alignment, Cross-Attention Fusion, Scene Understanding, Explainable Decision-Making.

Abstract

The traditional perception pipeline for autonomous driving uses solely vision-based systems to recognize and classify objects without the ability to understand natural language commands, provide explanations for their decisions, or comprehend unusual scenarios which lie out of their distribution. Vision-language models provide a richer understanding of the environment but even the adaptation of such systems to autonomous driving performs fusion of visual and language modalities separately, hence limiting the influence of each modality on another. In this paper, we propose Cross-Modal Intelligence Framework (CMIF) which learns the shared representation space of both modalities using contrastive pretraining, then uses the cross-modal fusion via a bi-directional cross-attention module to generate decisions. The drift detector will identify those examples for which there is less certainty about the degree of grounding and novelty and will perform targeted realignment about a buffer of such flagged corner cases, but not the entire retraining process. This framework was tested on a dataset modeled from the statistics of the nuScenes sensor data and BDD-X dataset of natural language descriptions of the driving scenarios, on four scenario classes by level of difficulty: normal driving, partial occlusion, corner-case events, and ambiguous language instructions. In comparison with the vision-only baseline and the late-fusion vision-language baseline models, the introduced CMIF achieves an 82.6 percent decision accuracy in case of ambiguous instructions against 45.2 percent and 62.9 percent respectively, with only slight latency overhead.

Downloads

Published

2026-06-14

How to Cite

Anitha, M., Sreedevi, D., Sowmya, D., Udayakumar, D., & Prasad, D. P. S. S. (2026). Cross-Modal Intelligence Framework for Vision-Language Integration in Autonomous Driving Applications. International Journal of Artificial Intelligence and Machine Learning, 6(5s), 589–596. Retrieved from https://www.svedbergopen.com/index.php/ijaiml/article/view/613