Logit-Contribution Scoring Identifies Non-Literal Retrieval Heads
By Aryo Pradipta Gema, Beatrice Alex, and Pasquale Minervini, July 1, 2026
In Mechanistic Interpretability Workshop at ICML 2026
Long-context language models often answer from the meaning of a relevant passage rather than copying its words. Existing head detectors largely reward literal token matches, so they can identify where a head attends while missing what its output-value circuit contributes to a non-literal answer.
Logit-Contribution Scoring (LOCOS) measures this write-side behaviour by projecting each head’s output-value contribution onto the answer token’s unembedding direction. It contrasts source positions inside and outside the relevant passage in one forward pass, producing a targeted ranking of retrieval heads.
Across Qwen3, Gemma-3, and OLMo-3.1 models, ablating highly ranked LOCOS heads disrupts non-literal retrieval more rapidly than attention-based detectors. The effects transfer to multi-hop and long-context tasks while leaving tested parametric recall and arithmetic behaviour near baseline, indicating that the selected heads are retrieval-specific.
Paper: https://arxiv.org/abs/2607.01002
Stay ahead with research-backed solutions
From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.
Book a Consultation