Can Retrieval Heads See Images? Multimodal Retrieval Heads in Long-Context Vision-Language Models

By Aaron Branson Cigres Li, Zhaowei Wang, Yu Zhao, Yiming Du, Haobo Li, Xiyu Ren, Ginny Wong, Simon See, Lishu Luo, Haodong Duan, Pasquale Minervini, and Yangqiu Song, May 26, 2026

In EMNLP 2026

Long-context vision-language models must find relevant evidence across interleaved text and images, but retrieval-head analyses designed for language models rely on token copying and do not transfer directly to visual evidence. This work introduces a multimodal detection method that scores attention from question tokens to either textual or visual evidence.

The analysis finds that multimodal retrieval heads are sparse, intrinsic to the models, and causally important. A small fraction of attention heads accounts for much of the positive retrieval score, while masking the highest-ranked heads sharply reduces performance on document and slide question-answering benchmarks compared with masking random heads.

These heads are partly shared across modalities but adapt to the context length and the modality of the evidence. Without additional training, their scores can also rank visually rich documents, connecting mechanistic analysis of vision-language models with practical multimodal retrieval.

Paper: https://arxiv.org/abs/2605.27243

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation