Rhetorical Effects on AI Scientific Review: A Critical Analysis
Summary of the Paper
This paper presents a controlled framework for studying how rhetorical rewriting affects AI reviewer judgments of scientific manuscripts. The researchers took anonymized papers, rewrote them along six dimensions (novelty framing, evidence emphasis, scope, clarity, structure, and caution) in either positive or negative directions, and then had five different AI reviewers evaluate both the original and rewritten manuscripts under standard and strict review protocols.
Key Findings
1. Rewriting has the largest effects on specific dimensions of presentation, and the direction of these effects depends on where the paper starts.
-
Evidence framing and novelty stance produced the largest shifts in Overall Assessment (OA): positive evidence framing raised OA by up to 0.93 (on a 10-point scale), while negative novelty framing lowered it.
-
Evidence reframing changed weak-accept probability by 13% on average.
-
Low-rated papers were more likely to rise, and high-rated papers more likely to fall, regardless of rewrite direction—suggesting constraints at the top and bottom of the rating scale.
-
Finding 2: More elaborate rewriting yields configuration-dependent and diminishing gains. Joint rewriting over six positive objectives produced mixed results (+2.49 under standard, +1.39 under strict Opus; +0.33 under standard, -1.09 under strict GPT). Recursive rewriting was neutral or negative in most conditionscompress.
-
Finding 3: stricter review lowers scores but preserves rewrite effects on human subjects. Both GPT-5.5 and Opus 4.8, especially stricter reviews.
2: Elaborate rewriting has diminishing effects
Joint rewriting that applies all six positive dimensions at once yields less than the sum of single-dimension effects asi gnal gains only under some conditions, recursive rewriting yields small or negative additional gains (Appendix E Figure 14), and reviewer-guided rewriting does no better than joint rewriting alone. These observations hold for almost every reviewer and condition we tested Our final enhancement attempts, in other words, procuse
Introduction
Presentation is part of peer review
Scientific manuscripts are expected to report results accurately, but also to frame them with the proper level of confidence accord- ing to the evidence, to position them appropriately with respect to the literature, and to communicate clearly to reviewers who do not have the same context as the authors. These rhetorical aspects (Fahnestock, 1998) are part of what makes a manuscript usable by peer reviewers amd sitummagem the
Related papers
- Shortcutting the Fix: Agentic Shortcutting in Software-Engineering Benchmarks
Agentic shortcutting—agents exploiting leaked solutions like upstream repos or Git history—inflates SWE benchmark scores by up to 82%, but a simple originality prompt cuts exploitation to under 11%.
- Frozen Judges, Moving Agents: Version-Dependent LLM-Judge Error and the Limits of Judge-Assisted Agent Evaluation
A fixed LLM judge produces version-dependent errors, invalidating agent comparisons and transported calibration, so release decisions require paired audits, not judge-only scores.
- Identical Runs, Different Results: Benchmarking AI Coding Agents on Open-Weight Models
Identical runs of the same AI coding agent vary more than differences between agent-model pairings, so best-of-three with compliance checking beats single-run benchmarking.