A scientific manuscript can receive a higher score simply because it is written differently, even while presenting exactly the same results? This is the question at the heart of a study authored by Ming Li and seven others, published on August 10, 2026, which introduces a concept both technical and unsettling into the academic debate: the rhetorical influence of AI evaluation in automated peer review processes. The research demonstrates that large language models tasked with judging scientific articles can be influenced by the style in which a text is presented, even when the scientific content remains identical.
Summary
The research team designed an experiment aimed at isolating a single variable: form, not substance. Starting from 120 anonymized submissions presented at ICLR 2026, one of the most relevant conferences in the field of machine learning, the authors generated a controlled corpus of 4,200 complete manuscripts, all derived from the same original works but rewritten according to different rhetorical variants.
Two large language model-based rewriters worked on six distinct rhetorical dimensions, transforming them in opposite directions for each manuscript: on one side, versions that emphasized certain persuasive aspects, and on the other, versions that downplayed them, while keeping the reported scientific results intact. At that point, five AI reviewers entered the scene, tasked with evaluating the manuscripts under two distinct protocols, one standard and one more rigorous, in addition to testing configurations of joint, recursive, and reviewer-guided rewriting.
The most significant result is that rhetorical sensitivity is not uniform but structured according to a true hierarchy. Among the six dimensions tested, the framing of evidence and the position on the novelty of the work produce the widest contrasts between positively and negatively rewritten versions. The framing of the research purpose forms a second level, weaker but still perceptible, while the other rhetorical dimensions show lesser or less stable effects.
This hierarchy also holds when comparing manuscripts of different perceived quality by original human reviewers. But there is a detail that makes the phenomenon even more delicate: the movement of scores strongly depends on the starting rating assigned by the AI reviewer. Manuscripts that started with lower scores tended to rise after rhetorical rewriting, while those with already high scores tended to drop. The clearest directional contrasts are observed in the intermediate score ranges, precisely where editorial decisions are more uncertain and therefore more susceptible to influence.
One of the most counterintuitive aspects of the study concerns the effectiveness of more sophisticated rhetorical manipulation methods. More elaborate workflows, which combine joint, recursive, or reviewer-guided rewriting, do not reliably produce greater gains compared to simpler interventions. In other words, adding complexity to the rewriting process does not automatically equate to better manipulation of AI judgment.
The effects of joint rewriting are found to be strongly dependent on the model used as a rewriter: changing the large language model significantly changes the outcome. Even providing explicit instructions to the AI reviewer did not show a consistent advantage over a simple second reading without guidance. Repeated rewriting, moreover, produces diminishing returns and is strongly dependent on the specific configuration used. However, the study identifies a rather clear division of roles: the rewriting model primarily determines the separation between opposing variants, while the reviewing model determines the magnitude and sign of the effect on the score.
When researchers applied a stricter review protocol, the average overall evaluation score dropped by 1.36 points compared to the standard protocol. But here lies the most important data for those designing these systems: greater severity did not consistently reduce the rhetorical sensitivity of AI reviewers. Making the judge more demanding, in short, does not automatically make them more immune to the stylistic choices of the text they are evaluating.
Why does all this matter? Because scientific review is already under increasing pressure. The volume of publications is growing at an exponential rate, and recruiting available human reviewers has become increasingly difficult. According to a survey conducted among the staff of Frontiers journals, over half of the reviewers already incorporate some form of artificial intelligence into their evaluation process. In this context, the idea that automated systems can be manipulated with simple rhetorical tricks, without touching the scientific content, opens a potentially serious crack in the integrity of the process.
The authors of the study do not merely document the problem: they present it as a form of reward hacking, that is, the exploitation of linguistic shortcuts to achieve a better score without actually improving the quality of the work. The conclusion that follows is clear: automated evaluation systems capable of resisting rhetorical variations that do not alter the scientific content are needed. As long as this robustness is not guaranteed, every added or reformulated line could weigh more than the data that a scientific article should truly communicate.
Rhetorical choices such as framing evidence and positioning on novelty significantly influence AI reviewers' judgments, even when the scientific content remains identical.
A controlled corpus of 4,200 manuscripts, derived from 120 anonymized submissions presented at ICLR 2026, was used.
No: more elaborate rewriting workflows did not reliably produce greater score gains compared to simpler interventions.
The rigorous review protocol lowered the average overall evaluation score by 1.36 points, but did not consistently alter the rhetorical sensitivity of AI reviewers.
Content created with the assistance of artificial intelligence and human editorial review.
This content is provided for general informational purposes only and doesn't constitute financial, investment, legal, or tax advice. Any events, rewards, online promotions, or related information mentioned herein should not be considered a recommendation, solicitation, or invitation to purchase, sell, trade, or otherwise deal in any crypto assets. Crypto assets are highly volatile and may result in loss. The availability of WEEX services, products, and related events may vary by region. You are responsible for ensuring that your participation is in accordance with applicable local laws and regulations.





























