Full text not available for this paper

False Frontiers: Diagnosing and Mitigating Co-Cheating in Self-Evolving Search Agents

Summary (Overview)

  • Identifies a novel failure mode called "co-cheating" in self-evolving search agents, where the proposer (question generator) and solver (answerer) increasingly agree on shared errors, causing internal reward to improve without corresponding gains in external correctness.
  • Introduces a post-hoc audit protocol using an LLM judge (gpt-6-astra/high) that constructs evidence-backed references from source documents, enabling scalable measurement of false-agreement mass (FF) as the fraction of pairs agreeing on the same incorrect answer.
  • Proposes two mitigation strategies: Multi-Sample Verification (MSV), which tests pseudo-label stability through source-aware and source-blind sampling, and CrossFit, which partitions source documents into folds and scores each fold's proposals with an auxiliary solver trained only on the complementary fold.
  • Demonstrates that CrossFit reduces false-agreement mass from 6.1%/8.8% to 3.0%/3.7% on Qwen3.5-4B/9B, while improving downstream performance by 8.8/8.4 points over standard coupled self-evolution and 8.7/7.8 points over Search-R1.
  • Shows through fixed-bank replay experiments that source-level exclusion of feedback ancestry, rather than evaluator duplication or task selection, is the operative mechanism responsible for improvement.

Introduction and Theoretical Foundation

Background

Search-augmented language models interleave reasoning with browser or search actions to gather evidence before answering. Most are trained on externally supplied questions and answer supervision. Self-evolving agents instead generate their own training experience through a proposer–solver architecture:

  1. A proposer turns source documents into questions and pseudo-labels
  2. Admitted pairs train a solver
  3. The solver's performance on new proposals determines the proposer reward

This closed loop creates an automated curriculum without a fixed human-authored training set.

The Co-Cheating Failure Mode

The loop makes agreement an endogenous proxy for correctness. The paper identifies a critical failure path:

"An incorrect pseudo-label can train the solver to repeat the same error on later questions from that source; rewarding this agreement then reinforces the error in the next-round curriculum."

The authors formalize this as co-cheating, measured as false-agreement mass (FF): the fraction of evaluated pairs that agree on the same incorrect answer. Key diagnostic metrics include:

  • TPT_P: adopted-label truth (correctness of the pseudo-label)
  • TST_S: solver-response truth (correctness of solver outputs)
  • AA: label–response match rate observed by the loop
  • FF: pairs matching the same incorrect answer (false agreement)
  • LL: correct solver responses denied credit by a wrong label (lost credit)

The theoretical foundation draws on prior work on reward hacking (Amodei et al., 2016), proxy reward optimization divergence (Gao et al., 2023), and pseudo-label confirmation bias (Arazo et al., 2020). The paper's unique contribution is identifying the additional return path from a pseudo-label-trained solver into task generation.

Methodology

Diagnostic Audit Protocol

The audit saves at every scheduled step:

  • The source document
  • The adopted pseudo-label
  • Five solver responses used for proposer reward

An LLM judge (gpt-6-astra/high) constructs an evidence-backed reference from the source and judges the saved outputs. Unsupported cases remain unresolved rather than receiving forced labels. The auditor never affects admission, model updates, or reward, providing an independent measurement.

Multi-Sample Verification (MSV)

MSV is an admission-time test of whether a proposed question admits a stable answer independently of the proposer's draft. Given source document xx, question qq, and model MM:

aisrc∼M(⋅∣x,q),aiblind∼M(⋅∣q),i∈{1,2,3}a^{\text{src}}_i \sim M(\cdot | x, q), \quad a^{\text{blind}}_i \sim M(\cdot | q), \quad i \in \{1, 2, 3\}

Let Maj\text{Maj} return an answer when at least two samples agree under answer matcher ≃\simeq, and ∅\emptyset otherwise. Defining yv=Maj(a1:3v)y^v = \text{Maj}(a^v_{1:3}) for v∈{src,blind}v \in \{\text{src}, \text{blind}\}:

IMSV=1[ysrc≠∅∧yblind≠∅∧ysrc≃yblind]I_{\text{MSV}} = \mathbb{1}[y^{\text{src}} \neq \emptyset \land y^{\text{blind}} \neq \emptyset \land y^{\text{src}} \simeq y^{\text{blind}}]

When IMSV=1I_{\text{MSV}} = 1, the compatible majority replaces the draft as the training label; otherwise the task is rejected.

CrossFit: Cross-Fitted Proposer Feedback

CrossFit changes only where the proposer obtains feedback:

  1. Source partitioning: Each source document is assigned once to fold 0 or fold 1 (at the source level to prevent related examples from the same document appearing on both sides)
  2. Auxiliary solvers: Two auxiliary feedback solvers train on their respective folds
  3. Cross-scoring: Questions from fold 0 are evaluated by the solver trained on fold 1, and vice versa
  4. Main solver: Trains on all admitted questions from both folds (unchanged)

The feedback rule remains the same as Dr. Zero:

RP(q)=f(∑j=151[zj≃y~]),zj∼Sr,1−h(⋅∣q)R_P(q) = f\left(\sum_{j=1}^{5} \mathbb{1}[z_j \simeq \tilde{y}]\right), \quad z_j \sim S_{r, 1-h}(\cdot | q)

where hh is the source fold, Sr,1−hS_{r, 1-h} is the auxiliary solver trained on the complementary fold, y~\tilde{y} is the adopted label, and f(k)=(5−k)/4f(k) = (5-k)/4 for 0<k<50 < k < 5 (zero otherwise).

Training Objective

The proposer and solver are updated with a sequence-normalized policy-gradient objective:

L(θ)=−1∣E∣∑e∈EA^e1∣Te∣∑t∈Telog⁡πθ(ye,t∣ye,<t,qe),A^e=re−μg(e)σg(e)+10−6\mathcal{L}(\theta) = -\frac{1}{|\mathcal{E}|} \sum_{e \in \mathcal{E}} \hat{A}_e \frac{1}{|T_e|} \sum_{t \in T_e} \log \pi_\theta(y_{e,t} | y_{e,<t}, q_e), \quad \hat{A}_e = \frac{r_e - \mu_{g(e)}}{\sigma_{g(e)} + 10^{-6}}

Empirical Validation / Results

Main Results

Table 1: Downstream search performance after three rounds of self-evolution (Cover-EM, 1,325-question set)

TreatmentNQTriviaQAPopQAHotpotQA2WikiMQAMuSiQueBamboogleAverage
Qwen3.5-4B
Base0.3800.6550.3150.3500.4650.1050.4160.384
Dr. Zero0.3900.6650.3250.3650.4850.1200.4480.400
MSV0.4000.6700.3400.3750.4800.1350.4480.407
CrossFit0.4550.7300.4150.4600.5750.2450.5360.488
MSV + CrossFit0.4700.7300.4100.4750.5750.2500.5280.491
Qwen3.5-9B
Base0.5050.7100.3750.3550.2950.1250.4960.409
Dr. Zero0.5150.7300.4000.3700.3200.1500.5120.428
MSV0.5250.7200.4000.3900.3250.1550.5360.436
CrossFit0.5800.7650.4550.4900.4250.2550.6160.512
MSV + CrossFit0.5700.7800.4700.4850.4150.2800.6080.515

Key findings:

  • CrossFit gains are largest on multi-hop tasks (averaging 10.0/10.9 points across HotpotQA, 2WikiMQA, MuSiQue, Bamboogle)
  • MSV alone improves only 0.7–0.8 points over Dr. Zero
  • MSV + CrossFit adds only 0.3 points over CrossFit alone

False-Agreement Reduction

Round-3 audit statistics (means of 43 round-3 step rates):

TreatmentTPT_PTST_SAAFFLL
Qwen3.5-4B
Dr. Zero0.7470.6690.7100.0610.020
MSV0.7470.7080.7450.0570.020
CrossFit0.8190.6860.6790.0300.038
MSV + CrossFit0.8430.7270.7080.0200.038
Qwen3.5-9B
Dr. Zero0.7370.6890.7520.0880.025
MSV0.7700.7270.7880.0720.011
CrossFit0.8510.7120.7090.0370.041
MSV + CrossFit0.8780.7540.7320.0170.039

Mechanism Ablations

The paper conducts fixed-bank replay experiments holding the bank, labels, answer matcher, and evaluation procedure fixed:

  • Same-source auxiliary solver: False-agreement mass of 0.064/0.087 (close to coupled control)
  • Full-data auxiliary: 0.058/0.069
  • Random question partitioning: Only modest reduction to 0.050/0.062
  • Source-ID split: Reduces false agreement to 0.004/0.001
  • Source-ID feedback on identical proposals: Reduces coupled false agreement from 0.058/0.073 to 0.004/0.001
  • Half-budget control: Matches full source-ID result (0.005/0.002), ruling out auxiliary optimization as explanation

Training Trajectory Analysis

Cross-fitted feedback reverses the divergence between agreement and truth:

  • Without cross-fitting, in-loop agreement rises above solver truth while false agreement accumulates
  • With cross-fitting, adopted-label truth rises, agreement remains at or below solver truth
  • By round 3, false-agreement mass falls below half of the coupled value

The downstream advantage of CrossFit over Dr. Zero grows from 4.2/4.3 points after round 2 to 8.8/8.4 points after round 3, showing the intervention changes what the loop learns across rounds.

Theoretical and Practical Implications

Theoretical Contributions

  1. Co-cheating as a distinct failure mode: The paper formalizes how agreement can improve because proposer and solver reinforce the same incorrect labels, distinguishing this from ordinary label noise or lost credit.

  2. Feedback ancestry matters: The key insight is that the training history of the evaluator matters as much as pseudo-label quality. A solver that trained on the same source-derived errors will systematically reproduce them as apparent progress.

  3. Cross-fitting principle: Borrowing from statistical cross-fitting (Chernozhukov et al., 2018), the paper demonstrates that source-excluded feedback prevents the direct self-reinforcing path without requiring a truth oracle.

Practical Implications

  1. Method ranking: CrossFit (0.488/0.512) > MSV + CrossFit (0.491/0.515) > MSV (0.407/0.436) > Dr. Zero (0.400/0.428) > Search-R1 (0.401/0.434) > Base (0.384/0.409)

  2. Cost considerations: MSV adds ~90% reserved budget and ~5x judge requests; CrossFit adds 72-79% (36-40% with halved auxiliary budget) with nearly identical replay false-agreement mass.

  3. Verification alone is insufficient: MSV improves pseudo-label reliability but does not remove the feedback dependence that produces co-cheating, as it doesn't prevent a later feedback solver from reusing labels derived from the evaluated source.

Conclusion

Co-cheating exposes a fundamental failure of self-evolution: agreement can improve because a proposer and solver reinforce the same incorrect labels. The paper's evidence-backed audit separates internal progress from correctness, and CrossFit addresses the feedback path by scoring each source with an auxiliary solver trained on the complementary fold while the main solver still learns from all admitted tasks.

Key results across Qwen3.5-4B and Qwen3.5-9B:

  • Reduces final-round false agreement from 6.1%/8.8% to 3.0%/3.7%
  • Improves seven-benchmark average Cover-EM over Dr. Zero by 8.8/8.4 points
  • Fixed-bank replay and evaluator controls support source ancestry as the operative distinction

Open limitations:

  • Shared pretraining errors and overlapping web evidence may still induce correlated mistakes
  • Lower false-agreement mass could partly result from rejecting difficult tasks
  • Extension to connected sources and end-to-end efficiency measurement remain essential next tests

Core principle: Reliable self-evolution requires auditing both feedback correctness and the training history of its evaluator.

Related papers