Inverse Scaling in Test-Time Compute

By Aryo Pradipta Gema, Alexander Hägele, Runjin Chen, Andy Arditi, Jacob Goldman Wetzler, Kit Fraser Taliente, Henry Sleight, Linda Petrini, Julian Michael, Beatrice Alex, Pasquale Minervini, Yanda Chen, Joe Benton, and Ethan Perez, July 19, 2025

In TMLR 2026

Allocating more test-time computation often improves large reasoning models, but the relationship is not universally monotonic. This work constructs tasks in which longer reasoning lowers accuracy, covering counting with distractors, regression with spurious features, constraint-tracking deductions, and advanced-AI-risk evaluations.

The experiments identify several mechanisms behind inverse scaling. Depending on the model family and task, extended reasoning can increase distraction, encourage overfitting to the problem framing, replace reasonable priors with spurious correlations, or make it harder to maintain focus through a complex deduction.

The results show why evaluation at a single reasoning budget can conceal important behaviour. Measuring performance across reasoning lengths is necessary both to characterise model capability and to detect undesirable patterns that additional computation may reinforce.

Paper: https://openreview.net/forum?id=NXgyHW1c7M

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation