Rhetorical Effects on AI Scientific Review: A Critical Analysis
Summary of the Paper
This paper presents a controlled framework for studying how rhetorical rewriting affects AI reviewer judgments of scientific manuscripts. The researchers took anonymized papers, rewrote them along six dimensions (novelty framing, evidence emphasis, scope, clarity, structure, and caution) in either positive or negative directions, and then had five different AI reviewers evaluate both the original and rewritten manuscripts under standard and strict review protocols.
Key Findings
1. Rewriting has the largest effects on specific dimensions of presentation, and the direction of these effects depends on where the paper starts.
-
Evidence framing and novelty stance produced the largest shifts in Overall Assessment (OA): positive evidence framing raised OA by up to 0.93 (on a 10-point scale), while negative novelty framing lowered it.
-
Evidence reframing changed weak-accept probability by 13% on average.
-
Low-rated papers were more likely to rise, and high-rated papers more likely to fall, regardless of rewrite direction—suggesting constraints at the top and bottom of the rating scale.
-
Finding 2: More elaborate rewriting yields configuration-dependent and diminishing gains. Joint rewriting over six positive objectives produced mixed results (+2.49 under standard, +1.39 under strict Opus; +0.33 under standard, -1.09 under strict GPT). Recursive rewriting was neutral or negative in most conditionscompress.
-
Finding 3: stricter review lowers scores but preserves rewrite effects on human subjects. Both GPT-5.5 and Opus 4.8, especially stricter reviews.
2: Elaborate rewriting has diminishing effects
Joint rewriting that applies all six positive dimensions at once yields less than the sum of single-dimension effects asi gnal gains only under some conditions, recursive rewriting yields small or negative additional gains (Appendix E Figure 14), and reviewer-guided rewriting does no better than joint rewriting alone. These observations hold for almost every reviewer and condition we tested Our final enhancement attempts, in other words, procuse
Introduction
Presentation is part of peer review
Scientific manuscripts are expected to report results accurately, but also to frame them with the proper level of confidence accord- ing to the evidence, to position them appropriately with respect to the literature, and to communicate clearly to reviewers who do not have the same context as the authors. These rhetorical aspects (Fahnestock, 1998) are part of what makes a manuscript usable by peer reviewers amd sitummagem the
Related papers
- Harness Continual Learning: Continual Adaptation Beyond Model Parameters
Harness Continual Learning enables frozen foundation models to accumulate capabilities by evolving prompts, memories, and tools around them, with guarded updates preventing harness-level forgetting.
- Rethinking Self-Evolving Agents: Do We Still Need Prescribed Optimization Pipelines?
Frontier optimizers can compose task-specific improvement strategies online without prescribed pipelines, matching or beating them on 12 of 14 settings while using a third of the compute.
- More Convincing, Not More Correct: Self-Play Reward Hacking of Reference-Free LLM Judges
Self-play training against reference-free LLM judges inflates pass rates without improving true accuracy, creating a 0.74 judge–truth gap; forcing judges to commit their own answer first collapses false positives from 0.719 to 0.012.