SCOPE: Self-Play via Co-Evolving Policies for Open-Ended Tasks

By Wai Chung Kwan, Aryo Pradipta Gema, Joshua Ong Jun Leang, and Pasquale Minervini, May 29, 2026

In arXiv 2026

Most language-model self-play methods depend on answers that can be checked with fixed rules. Open-ended tasks instead tend to require curated prompts or stronger external models to judge responses.

SCOPE removes these dependencies by co-evolving two policies: a Challenger that creates document-grounded tasks and a Solver that answers them through multi-turn retrieval. A frozen copy of the initial model generates task-specific rubrics and evaluates the answers.

Across several instruction-tuned models, SCOPE improves performance on open-ended benchmarks and transfers to held-out short-form question answering, showing that self-play can develop both retrieval and synthesis capabilities without externally supplied training tasks.

Paper: https://arxiv.org/abs/2605.31433

Stay ahead with research-backed solutions

From papers to production, we translate cutting-edge AI research into practical systems that give your business a competitive edge.

Book a Consultation