OpenSIR: Open-Ended Self-Improving Reasoner
By Wai Chung Kwan, Joshua Ong Jun Leang, Pavlos Vougiouklis, Jeff Z. Pan, Marco Valentino, and Pasquale Minervini, November 1, 2025
In arXiv 2025
OpenSIR is a self-play framework in which one language model alternates between teacher and student roles, generating and solving new problems without annotated datasets or external verifiers.
Starting from a single seed problem, the framework rewards conceptual diversity and calibrates difficulty so that generated problems remain both novel and learnable. This prevents self-play from repeatedly producing tasks around familiar concepts.
OpenSIR improves instruction and reasoning models across mathematical benchmarks and transfers beyond its self-generated mathematics training to broader reasoning tasks.
Paper: https://arxiv.org/abs/2511.00602
Stay ahead with research-backed solutions
From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.
Book a Consultation