OpenSIR: Open-Ended Self-Improving Reasoner

By Wai Chung Kwan, Joshua Ong Jun Leang, Pavlos Vougiouklis, Jeff Z. Pan, Marco Valentino, and Pasquale Minervini, November 1, 2025

In arXiv 2025

OpenSIR is a self-play framework in which one language model alternates between teacher and student roles, generating and solving new problems without annotated datasets or external verifiers.

Starting from a single seed problem, the framework rewards conceptual diversity and calibrates difficulty so that generated problems remain both novel and learnable. This prevents self-play from repeatedly producing tasks around familiar concepts.

OpenSIR improves instruction and reasoning models across mathematical benchmarks and transfers beyond its self-generated mathematics training to broader reasoning tasks.

Paper: https://arxiv.org/abs/2511.00602

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation