Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads

By Aryo Pradipta Gema, Beatrice Alex, and Pasquale Minervini, July 1, 2026

In Mechanistic Interpretability Workshop at ICML 2026

Long-context language models often answer from the meaning of a relevant passage rather than copying its words. Existing head detectors largely reward literal token matches, so they can identify where a head attends while missing what its output-value circuit contributes to a non-literal answer.

Logit-Contribution Scoring (LOCOS) measures this write-side behaviour by projecting each head’s output-value contribution onto the answer token’s unembedding direction. It contrasts source positions inside and outside the relevant passage in one forward pass, producing a targeted ranking of retrieval heads.

Across Qwen3, Gemma-3, and OLMo-3.1 models, ablating highly ranked LOCOS heads disrupts non-literal retrieval more rapidly than attention-based detectors. The effects transfer to multi-hop and long-context tasks while leaving tested parametric recall and arithmetic behaviour near baseline, indicating that the selected heads are retrieval-specific.

Paper: https://arxiv.org/abs/2607.01002

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation