Hera-Vlm: Hallucination-Resistant Evidence Retrieval and Attribution for Vision-Language Models
Keywords:
Vision-Language Models, Multimodal Reasoning, Hallucination Detection, Evidence Retrieval, Evidence Attribution, Visual Grounding, Qwen2-VL.Abstract
Despite being able to reason about images and answer questions in natural language, Vision-Language Models (VLMs) can hallucinate information not present in the source material, leading to unreliable responses from multimodal language models, especially when the answer requires multiple reasoning steps. In this paper, we present HERA-VLM (Hallucination-resistant Evidence Retrieval and Attribution for Vision-Language models), a novel framework for enhanced multimodal response generation with evidence retrieval and step-level verification. Our approach takes as input both the image and question, and outputs both the answer and reasoning trace produced by the Qwen2-VL-7B-Instruct model. We decompose the reasoning trace into individual reasoning steps. For each, we retrieve supporting evidence. For claims that are image-based, we extract region-based visual evidence, but for text-based claims, we optionally perform a web search to retrieve additional supporting information. We then use a cross-encoder model to evaluate how well each piece of evidence supports the corresponding claim and obtain a confidence score. By analyzing the verification results for each step, it becomes possible to identify which claims are well-supported, questionable, or hallucinated. Evidence can be cited explicitly for each reasoning step to improve answer transparency. Questionable reasoning steps can be revised to produce improved grounded responses. The implementation described in this paper is an inference-time only prototype that does not require training specialized VLMs by utilizing off-the-shelf pretrained models.





