PiCSAR: Probabilistic Confidence Selection and Ranking for Reasoning Chains
By Joshua Ong Jun Leang, Zheng Zhao, Aryo Pradipta Gema, Sohee Yang, Wai Chung Kwan, Xuanli He, Wenda Li, Pasquale Minervini, Eleonora Giunchiglia, and Shay B. Cohen, July 2, 2026
In Findings of ACL 2026
Best-of-n inference can improve reasoning by sampling several candidate solutions, but it still needs a reliable way to select a correct chain without access to the answer. PiCSAR provides a training-free score based on the joint log-likelihood of each candidate’s reasoning path and final answer.
This joint score decomposes naturally into reasoning confidence and answer confidence, allowing the selector to account for the model’s certainty throughout a derivation rather than judging only its conclusion. The method can be applied directly to existing language and reasoning models.
Across the evaluated reasoning benchmarks, PiCSAR improves selection accuracy and often outperforms baselines with at least two times fewer samples. Analysis also finds that correct chains tend to receive higher confidence for both their reasoning and their answers.
Paper: https://aclanthology.org/2026.findings-acl.1577/
Stay ahead with research-backed solutions
From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.
Book a Consultation