Mitigating multimodal hallucinations through visual attention tracing and origin-point regeneration

B Bohan Li H Haiyang Yu Y Yishan Han J Julong Ren M Mengyi Dai S Shumei Bao

Abstract

Abstract Multimodal large language models (MLLMs) exhibit impressive prowess in vision-language understanding, but their utility is often compromised by hallucinations—instances where generated narratives diverge significantly from visual evidence. Current remedial strategies largely struggle to pinpoint the genesis of these errors, relying either on resource-intensive retraining or indiscriminate global penalties that fail to address the specific locus of the discrepancy. Addressing this limitation, we introduce hallucination backtracking (HB), a training-free decoding framework designed to effectively detect and mitigate errors by monitoring visual attention dynamics during generation. This approach is grounded in the observation that hallucinations are not random; rather, they stem from specific pivotal tokens where the model’s focus precipitously shifts from image features to its own prior textual generations. By quantifying this drift through a novel visual attention score (VAS), our origin-point detection mechanism successfully isolates the source of errors, achieving a 41.8% exact match and 84.1% before-first accuracy in localization. Once an attentional anomaly is detected, the system autonomously backtracks to the divergence point, triggering a regeneration process reinforced by stricter visual grounding constraints. Rigorous evaluations across diverse architectures—LLaVA-1.5, InstructBLIP, MiniGPT-4, and Shikra—confirm that HB consistently surpasses state-of-the-art baselines; notably, on LLaVA-1.5, our method elevates the F1 score on the POPE benchmark to 91.4% while reducing the CHAIR $$_S$$ metric to 40.2%, yielding improvements of 1.5 and 4.4 points over OPERA, respectively, though a residual false negative rate of 15.9% indicates that inference-driven hallucinations remain an open challenge. Beyond quantitative gains, we provide a granular dissection of the hallucination phenomenon through analyses of VAS trajectory patterns and failure modes, ultimately advocating for precise localization and targeted correction as a promising paradigm for reliable multimodal generation.

Article Details

Volume / Issue Vol. 1, Issue 1
Published June 04, 2026
ISSN 2045-2322
Publisher Nature Portfolio

Journal Info

Scientific Reports

Nature Portfolio

ISSN: 2045-2322 Open Access Life Sciences

Authors (6)

B

Bohan Li

H

Haiyang Yu

Y

Yishan Han

J

Julong Ren

M

Mengyi Dai

S

Shumei Bao