Domain-Specialized Multi-Agent Collaboration with Consensus-Aware Fusion for Remote Sensing Image Captioning
Keywords:
Remote Sensing (RS), Satellite Image Processing, RS Caption generation, Land Cover ClassificationAbstract
Now a days Remote sensing image captioning (RSIC) helps a lot in understanding complex satellite images by generating human-understandable textual descriptions. These textual descriptions useful in many domains such as land-use analysis, disaster management, urban planning, and environmental monitoring. For generating such descriptions, most of the current methods use single-agent encoder-decoder frameworks, which perform well to some extent but generally have trouble with domain generalization, semantic richness, and robustness in the varied and heterogeneous environments shown in remote sensing images. To overcome these restrictions, we presented MAC-CapNet (Multi-Agent Collaborative Captioning Network), an innovative architecture that utilizes the collaboration of many specialized agents, and each agent trained to concentrate on a certain land-use domain (e.g., urban, forest, agricultural, water bodies, barren land ). These agents independently gives us candidate captions and confidence scores, which are then combined using a dynamically weighted meta-decoder that takes into account both contextual relevance and agent trustworthiness. The final caption is generated through a context-aware gating technique. The improvements we made include: (1) a multi-agent architecture designed specifically for RSIC; (2) a unique meta-decoder that dynamically combines agent predictions; and (3) numerous experiments evidence of higher performance on standard datasets. Evaluations carried out with RSICD and UCM-Captions demonstrate considerable gains in BLEU-4, METEOR, and CIDEr scores compared to cutting-edge approaches. The proposed MAC-CapNet was evaluated on widely used benchmark datasets for remote sensing image captioning. Across all experiments, it achieved better overall performance than the baseline models on BLEU-4, METEOR, ROUGE-L, CIDEr, and SPICE. We also carried out domain-wise comparisons, robustness experiments, scalability analysis, and statistical evaluation. Together, these results show that the framework performs reliably under different experimental settings and supports the design choices introduced in this work.





