Full text not available for this paper
False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents
Summary (Overview)
- Identifies a novel failure mode called "co-cheating" in self-evolving search agents, where the proposer (question generator) and solver (answerer) increasingly agree on shared errors, causing internal reward to improve without corresponding gains in external correctness.
- Introduces a post-hoc audit protocol using an LLM judge (gpt-6-astra/high) that constructs evidence-backed references from source documents, enabling scalable measurement of false-agreement mass () as the fraction of pairs agreeing on the same incorrect answer.
- Proposes two mitigation strategies: Multi-Sample Verification (MSV), which tests pseudo-label stability through source-aware and source-blind sampling, and CrossFit, which partitions source documents into folds and scores each fold's proposals with an auxiliary solver trained only on the complementary fold.
- Demonstrates that CrossFit reduces false-agreement mass from 6.1%/8.8% to 3.0%/3.7% on Qwen3.5-4B/9B, while improving downstream performance by 8.8/8.4 points over standard coupled self-evolution and 8.7/7.8 points over Search-R1.
- Shows through fixed-bank replay experiments that source-level exclusion of feedback ancestry, rather than evaluator duplication or task selection, is the operative mechanism responsible for improvement.
Introduction and Theoretical Foundation
Background
Search-augmented language models interleave reasoning with browser or search actions to gather evidence before answering. Most are trained on externally supplied questions and answer supervision. Self-evolving agents instead generate their own training experience through a proposer–solver architecture:
- A proposer turns source documents into questions and pseudo-labels
- Admitted pairs train a solver
- The solver's performance on new proposals determines the proposer reward
This closed loop creates an automated curriculum without a fixed human-authored training set.
The Co-Cheating Failure Mode
The loop makes agreement an endogenous proxy for correctness. The paper identifies a critical failure path:
"An incorrect pseudo-label can train the solver to repeat the same error on later questions from that source; rewarding this agreement then reinforces the error in the next-round curriculum."
The authors formalize this as co-cheating, measured as false-agreement mass (): the fraction of evaluated pairs that agree on the same incorrect answer. Key diagnostic metrics include:
- : adopted-label truth (correctness of the pseudo-label)
- : solver-response truth (correctness of solver outputs)
- : label–response match rate observed by the loop
- : pairs matching the same incorrect answer (false agreement)
- : correct solver responses denied credit by a wrong label (lost credit)
The theoretical foundation draws on prior work on reward hacking (Amodei et al., 2016), proxy reward optimization divergence (Gao et al., 2023), and pseudo-label confirmation bias (Arazo et al., 2020). The paper's unique contribution is identifying the additional return path from a pseudo-label-trained solver into task generation.
Methodology
Diagnostic Audit Protocol
The audit saves at every scheduled step:
- The source document
- The adopted pseudo-label
- Five solver responses used for proposer reward
An LLM judge (gpt-6-astra/high) constructs an evidence-backed reference from the source and judges the saved outputs. Unsupported cases remain unresolved rather than receiving forced labels. The auditor never affects admission, model updates, or reward, providing an independent measurement.
Multi-Sample Verification (MSV)
MSV is an admission-time test of whether a proposed question admits a stable answer independently of the proposer's draft. Given source document , question , and model :
Let return an answer when at least two samples agree under answer matcher , and otherwise. Defining for :
When , the compatible majority replaces the draft as the training label; otherwise the task is rejected.
CrossFit: Cross-Fitted Proposer Feedback
CrossFit changes only where the proposer obtains feedback:
- Source partitioning: Each source document is assigned once to fold 0 or fold 1 (at the source level to prevent related examples from the same document appearing on both sides)
- Auxiliary solvers: Two auxiliary feedback solvers train on their respective folds
- Cross-scoring: Questions from fold 0 are evaluated by the solver trained on fold 1, and vice versa
- Main solver: Trains on all admitted questions from both folds (unchanged)
The feedback rule remains the same as Dr. Zero:
where is the source fold, is the auxiliary solver trained on the complementary fold, is the adopted label, and for (zero otherwise).
Training Objective
The proposer and solver are updated with a sequence-normalized policy-gradient objective:
Empirical Validation / Results
Main Results
Table 1: Downstream search performance after three rounds of self-evolution (Cover-EM, 1,325-question set)
| Treatment | NQ | TriviaQA | PopQA | HotpotQA | 2WikiMQA | MuSiQue | Bamboogle | Average |
|---|---|---|---|---|---|---|---|---|
| Qwen3.5-4B | ||||||||
| Base | 0.380 | 0.655 | 0.315 | 0.350 | 0.465 | 0.105 | 0.416 | 0.384 |
| Dr. Zero | 0.390 | 0.665 | 0.325 | 0.365 | 0.485 | 0.120 | 0.448 | 0.400 |
| MSV | 0.400 | 0.670 | 0.340 | 0.375 | 0.480 | 0.135 | 0.448 | 0.407 |
| CrossFit | 0.455 | 0.730 | 0.415 | 0.460 | 0.575 | 0.245 | 0.536 | 0.488 |
| MSV + CrossFit | 0.470 | 0.730 | 0.410 | 0.475 | 0.575 | 0.250 | 0.528 | 0.491 |
| Qwen3.5-9B | ||||||||
| Base | 0.505 | 0.710 | 0.375 | 0.355 | 0.295 | 0.125 | 0.496 | 0.409 |
| Dr. Zero | 0.515 | 0.730 | 0.400 | 0.370 | 0.320 | 0.150 | 0.512 | 0.428 |
| MSV | 0.525 | 0.720 | 0.400 | 0.390 | 0.325 | 0.155 | 0.536 | 0.436 |
| CrossFit | 0.580 | 0.765 | 0.455 | 0.490 | 0.425 | 0.255 | 0.616 | 0.512 |
| MSV + CrossFit | 0.570 | 0.780 | 0.470 | 0.485 | 0.415 | 0.280 | 0.608 | 0.515 |
Key findings:
- CrossFit gains are largest on multi-hop tasks (averaging 10.0/10.9 points across HotpotQA, 2WikiMQA, MuSiQue, Bamboogle)
- MSV alone improves only 0.7–0.8 points over Dr. Zero
- MSV + CrossFit adds only 0.3 points over CrossFit alone
False-Agreement Reduction
Round-3 audit statistics (means of 43 round-3 step rates):
| Treatment | |||||
|---|---|---|---|---|---|
| Qwen3.5-4B | |||||
| Dr. Zero | 0.747 | 0.669 | 0.710 | 0.061 | 0.020 |
| MSV | 0.747 | 0.708 | 0.745 | 0.057 | 0.020 |
| CrossFit | 0.819 | 0.686 | 0.679 | 0.030 | 0.038 |
| MSV + CrossFit | 0.843 | 0.727 | 0.708 | 0.020 | 0.038 |
| Qwen3.5-9B | |||||
| Dr. Zero | 0.737 | 0.689 | 0.752 | 0.088 | 0.025 |
| MSV | 0.770 | 0.727 | 0.788 | 0.072 | 0.011 |
| CrossFit | 0.851 | 0.712 | 0.709 | 0.037 | 0.041 |
| MSV + CrossFit | 0.878 | 0.754 | 0.732 | 0.017 | 0.039 |
Mechanism Ablations
The paper conducts fixed-bank replay experiments holding the bank, labels, answer matcher, and evaluation procedure fixed:
- Same-source auxiliary solver: False-agreement mass of 0.064/0.087 (close to coupled control)
- Full-data auxiliary: 0.058/0.069
- Random question partitioning: Only modest reduction to 0.050/0.062
- Source-ID split: Reduces false agreement to 0.004/0.001
- Source-ID feedback on identical proposals: Reduces coupled false agreement from 0.058/0.073 to 0.004/0.001
- Half-budget control: Matches full source-ID result (0.005/0.002), ruling out auxiliary optimization as explanation
Training Trajectory Analysis
Cross-fitted feedback reverses the divergence between agreement and truth:
- Without cross-fitting, in-loop agreement rises above solver truth while false agreement accumulates
- With cross-fitting, adopted-label truth rises, agreement remains at or below solver truth
- By round 3, false-agreement mass falls below half of the coupled value
The downstream advantage of CrossFit over Dr. Zero grows from 4.2/4.3 points after round 2 to 8.8/8.4 points after round 3, showing the intervention changes what the loop learns across rounds.
Theoretical and Practical Implications
Theoretical Contributions
-
Co-cheating as a distinct failure mode: The paper formalizes how agreement can improve because proposer and solver reinforce the same incorrect labels, distinguishing this from ordinary label noise or lost credit.
-
Feedback ancestry matters: The key insight is that the training history of the evaluator matters as much as pseudo-label quality. A solver that trained on the same source-derived errors will systematically reproduce them as apparent progress.
-
Cross-fitting principle: Borrowing from statistical cross-fitting (Chernozhukov et al., 2018), the paper demonstrates that source-excluded feedback prevents the direct self-reinforcing path without requiring a truth oracle.
Practical Implications
-
Method ranking: CrossFit (0.488/0.512) > MSV + CrossFit (0.491/0.515) > MSV (0.407/0.436) > Dr. Zero (0.400/0.428) > Search-R1 (0.401/0.434) > Base (0.384/0.409)
-
Cost considerations: MSV adds ~90% reserved budget and ~5x judge requests; CrossFit adds 72-79% (36-40% with halved auxiliary budget) with nearly identical replay false-agreement mass.
-
Verification alone is insufficient: MSV improves pseudo-label reliability but does not remove the feedback dependence that produces co-cheating, as it doesn't prevent a later feedback solver from reusing labels derived from the evaluated source.
Conclusion
Co-cheating exposes a fundamental failure of self-evolution: agreement can improve because a proposer and solver reinforce the same incorrect labels. The paper's evidence-backed audit separates internal progress from correctness, and CrossFit addresses the feedback path by scoring each source with an auxiliary solver trained on the complementary fold while the main solver still learns from all admitted tasks.
Key results across Qwen3.5-4B and Qwen3.5-9B:
- Reduces final-round false agreement from 6.1%/8.8% to 3.0%/3.7%
- Improves seven-benchmark average Cover-EM over Dr. Zero by 8.8/8.4 points
- Fixed-bank replay and evaluator controls support source ancestry as the operative distinction
Open limitations:
- Shared pretraining errors and overlapping web evidence may still induce correlated mistakes
- Lower false-agreement mass could partly result from rejecting difficult tasks
- Extension to connected sources and end-to-end efficiency measurement remain essential next tests
Core principle: Reliable self-evolution requires auditing both feedback correctness and the training history of its evaluator.
Related papers
- PanoVLN: Towards Effective Panoramic Vision-and-Language Navigation
PanoVLN achieves state-of-the-art vision-and-language navigation by pairing panoramic 360-degree inputs with longer action horizons, confidence-guided execution, and geometry-aware visual fusion.
- Omni-IO Skills: Harnessing Your Agent Omni-Native
Omni-IO Skills is a plug-and-play harness that makes general-purpose agents omni-native without retraining, achieving 100% input-support and large quality gains on UniM-90.
- In-Context Learning for Robots: Methods and Applications
In-context learning lets robots adapt behavior from demonstrations and corrections without parameter updates, with success depending on preserving task-relevant distinctions from evidence to execution.