Summary of "From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement"
Summary (Overview)
- New paradigm (RLSVR): Extends RLVR to open-ended tasks by transforming the task into a proxy environment where rewards become verifiable by construction, inspired by self-supervised learning’s pretext-task principle.
- SpyRL instantiation: An information-asymmetric self-play game based on “Who Is the Spy?” where a predetermined spy identity turns output-quality assessment into a verifiable identity-recognition problem.
- State-of-the-art results: SpyRL outperforms existing self-improvement methods (R-Zero, Absolute Zero) on non-verifiable tasks (summarization, creative writing) and yields consistent gains on verifiable mathematical reasoning, with vote-based rewards closely aligned with human and LLM quality judgments.
- Verifiability is engineered: The work demonstrates that verifiability need not be an intrinsic property of a task—it can be created through task transformation, opening a path toward scalable, verifier-free self-improvement for general open-ended capabilities.
Introduction and Theoretical Foundation
Reinforcement Learning with Verifiable Rewards (RLVR) has driven progress in reasoning models (e.g., OpenAI o1, DeepSeek-R1) by enabling large-scale optimization with exact, rule-based reward signals. However, RLVR is limited to domains where correctness can be deterministically verified (e.g., math, coding). Open-ended tasks (creative writing, summarization) rely on human preferences, learned reward models, or LLM-as-a-Judge, introducing evaluation bias, judge capability bottlenecks, and additional inference costs.
The authors draw inspiration from self-supervised learning (SSL), which constructs pretext tasks whose supervisory signals are derived automatically from the data itself (e.g., masked language modeling, contrastive learning). SSL shows that when a proxy objective generates supervision automatically while preserving substantial capability overlap with the target task, learning can proceed without task-specific human annotation.
RLSVR extends this task-transformation principle to RLVR. Instead of approximating an unverifiable quality function with an external evaluator, RLSVR transforms the original open-ended task into a proxy environment that:
- Injects a latent variable (known only to the environment).
- Conditions task execution on observations derived from .
- Defines a verifiable interaction (e.g., identity inference) whose outcome can be checked exactly against .
- Computes a rule-based reward from the interaction outcome.
The reward is called self-verifiable because it is deterministic, rule-based, and requires no human annotation, learned reward model, or external judge. RLSVR is thus "self-supervised learning for RLVR": the transformation corresponds to the pretext task, to the automatically generated label, and standard RLVR optimization (e.g., GRPO) applies directly to .
Methodology
SpyRL is a concrete instantiation of RLSVR using information-asymmetric self-play. The framework consists of two stages per training epoch (see Algorithm 1):
1. Information-Asymmetric Performing Stage
- Sample an input and a spy index .
- Assign observations: civilians receive full ; the spy receives a degraded version (e.g., continuous span masking of 20%–40% of the text).
- Each player generates an output for the target task (summarization, creative writing, or math problem construction/solution).
2. Detection Stage with Verifiable Rewards
- All outputs are revealed.
- Each player votes on the spy identity: .
- Detection reward: (deterministically verifiable).
- Group-based advantage (GRPO-style): , normalized within the group.
3. Two-Stage Coupled Optimization
- Performing reward (zero-sum between spy and civilians): where is votes received by the spy, by civilian , the average civilian votes, and control competition strength.
- Role-Advantage Estimation (RAE): Calibrates role biases by subtracting role-specific baselines, preventing the optimizer from conflating the spy’s information disadvantage with poor policy quality.
- Alternating optimization: The performer and detector are updated alternately to avoid policy stagnation and maintain stable learning pressure.
- KL-regularized GRPO objectives:
The degradation operator is domain-specific: continuous span masking (20% for summarization/writing, 40% for math reasoning). All experiments use group size (default), batch size 1024, 100 epochs, and Qwen3-4B/8B backbones.
Empirical Validation / Results
Main Results on Non-Verifiable and Verifiable Tasks
Table 1: Summarization benchmarks (ROUGE-L and GPT-4o A/B win rates of SpyRL vs. each baseline)
| Method | GovReport | Multi_News | QmSum | VcSum | SamSum |
|---|---|---|---|---|---|
| Qwen3-4B | 30.2 / 74.6% | 23.1 / 80.2% | 21.3 / 68.4% | 15.1 / 70.2% | 43.2 / 76.2% |
| + R-Zero | 32.1 / 72.1% | 22.4 / 86.6% | 21.5 / 68.2% | 15.6 / 74.6% | 42.8 / 81.2% |
| + Absolute Zero | 33.2 / 68.5% | 25.2 / 78.4% | 22.7 / 64.8% | 18.3 / 66.8% | 46.1 / 73.4% |
| + SpyRL | 36.7 / – | 26.4 / – | 25.3 / – | 19.1 / – | 48.2 / – |
| Qwen3-8B | 29.0 / 78.2% | 23.1 / 68.5% | 19.2 / 78.2% | 14.9 / 72.5% | 44.3 / 79.5% |
| + R-Zero | 29.4 / 80.3% | 22.2 / 67.5% | 18.8 / 83.9% | 14.9 / 70.4% | 44.8 / 80.0% |
| + Absolute Zero | 32.5 / 74.2% | 23.2 / 68.3% | 19.1 / 78.8% | 15.8 / 68.2% | 46.2 / 70.4% |
| + SpyRL | 34.1 / – | 25.8 / – | 23.2 / – | 19.1 / – | 48.5 / – |
Table 2: Creative writing (GPT-4o pairwise win rates; SpyRL vs. each baseline)
| Method | WritingPrompt (Overall) | WritingBench (Overall) |
|---|---|---|
| Qwen3-4B | 81.3% | 75.1% |
| R-Zero | 78.9% | 75.0% |
| Absolute Zero | 75.6% | 71.1% |
| SpyRL (Qwen3-8B vs. Qwen3-8B) | 76.5% | 78.1% |
Table 3: Mathematical reasoning (accuracy %)
| Method | GSM8K | Math500 | AIME 24 | AIME 25 | Minerva | MMLU-Pro | GPQA-D |
|---|---|---|---|---|---|---|---|
| Qwen3-4B | 84.5 | 68.2 | 10.3 | 6.7 | 42.3 | 51.6 | 26.3 |
| + R-Zero | 88.7 | 72.8 | 10.3 | 6.7 | 47.1 | 52.8 | 27.8 |
| + Absolute Zero | 89.3 | 76.2 | 12.2 | 13.4 | 41.9 | 52.6 | 35.3 |
| + SpyRL | 93.4 | 79.5 | 13.3 | 20.0 | 47.8 | 57.4 | 41.3 |
| Qwen3-8B | 91.8 | 74.2 | 15.3 | 12.1 | 49.3 | 58.1 | 33.3 |
| + R-Zero | 92.1 | 78.4 | 15.3 | 14.2 | 52.5 | 61.7 | 34.3 |
| + Absolute Zero | 92.0 | 76.6 | 18.4 | 18.2 | 52.9 | 62.5 | 36.8 |
| + SpyRL | 93.5 | 81.2 | 20.0 | 23.3 | 56.3 | 63.1 | 39.8 |
Additional Validations
- Human evaluation (Table 4): SpyRL achieves 80.0% overall win rate vs. Qwen3-4B on WritingPrompt, consistently higher across novelty, emotion, coherence, and consistency.
- Rubric-as-reward comparison (Table 5): SpyRL outperforms Qwen3.5-27B-RaR (59.3% overall win) and is competitive with GPT-4o-RaR while incurring zero external verifier cost.
- Domain-specific summarization (Table 6): Trained on PubMed, SpyRL raises ROUGE-L on arXiv, PubMed, and BillSum by an average of 4.9 points.
- Cross-task transfer (Table 7): Summarization and creative writing transfer positively in both directions; math reasoning does not transfer to writing tasks.
- Ablation (Table 8): Full two-stage optimization is crucial; removing it leads to rapid plateauing.
- Group size (Figure 5): provides the best trade-off (mean gain from 5.5 to 9.3 over base); further scaling shows diminishing returns.
- Role-Advantage Estimation (Table 9): Removing RAE degrades the average from 50.4 to 37.5, actively harming the model.
- Degradation operator sensitivity (Table 10): 20% vs. 40% masking ratio yields nearly indistinguishable results, showing robustness.
Figure 4 (vote-quality correlation): A positive correlation between the number of suspicion votes received and GPT-4o rank (lower quality → more votes) confirms that the vote-based reward aligns with actual task performance without external verifiers.
Theoretical and Practical Implications
- Theoretical: RLSVR unifies two previously separate approaches—RLVR (exact verification) and SSL (label-by-construction)—showing that verifiability can be engineered through task transformation rather than being an intrinsic property of a task. The latent variable plays the role of automatically generated labels, and the proxy environment provides a rule-based reward that is exact and unbiased.
- Practical: SpyRL provides a scalable, verifier-free self-improvement pipeline for open-ended tasks. It eliminates the need for expensive human annotations, learned reward models, or LLM judges, reducing both cost and evaluation bias. The method is computationally efficient (no external verifier calls) and generalizes across domains (summarization, creative writing, reasoning). The cross-task transfer results suggest that training on one open-ended task can improve performance on related tasks.
- Limitations: The degradation operator must be designed per task, though experiments show it is robust to masking ratio. The framework currently requires multiple agents (default ), which increases inference cost per training step. The spy mechanism relies on the assumption that information deficit leads to detectably lower quality; tasks where this is not true (e.g., tasks where the spy can compensate) may require alternative transformations.
Conclusion
The paper proposes RLSVR, a paradigm that extends RLVR to open-ended tasks by transforming them into proxy environments with self-verifiable rewards. The instantiation SpyRL uses information-asymmetric self-play to turn output-quality assessment into a verifiable identity-recognition problem. Extensive experiments on summarization, creative writing, and mathematical reasoning show that SpyRL outperforms existing self-improvement methods (R-Zero, Absolute Zero) and even yields gains on already-verifiable reasoning tasks. The vote-based rewards align closely with human and LLM quality judgments. The key takeaway is that verifiability can be engineered through task transformation, opening a path toward scalable, verifier-free self-improvement for general open-ended capabilities. Models and code are released at https://github.com/wangqinsi1/RLSVR/tree/SpyRL.
Related papers
- SOLAR: A State-Driven Online Learning Rate Scheduler for LLM Pretraining
SOLAR uses reinforcement learning to learn bounded residual corrections to a base learning-rate schedule, improving LLM pretraining perplexity across dense and MoE models up to 3B parameters.
- Just Let Linear States Forget the Distant Past: Prefix Caching via Suffix Replay for Hybrid LLMs
SuffixReplay enables fine-grained prefix caching in hybrid LLMs by reconstructing linear-attention states from sparse anchors, cutting storage by 2x and TTFT by up to 70%.
- Learning Meta-Skills for Agent Harness Design in Test-Time AI4AI
Learning reusable meta-skills for environment design improves AI test-time performance by 8.95 points over no-skill construction, enabling fixed-weight self-improvement.