AI sycophancy is systematically dismantling academic evaluation

AI sycophancy is systematically dismantling academic evaluation. Ke-ke Shang, Computational Communication Collaboratory, Nanjing University 
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Two years ago, a paper I co-authored "died" in peer review. What killed it was not bad science. It was three references that did not exist, fabricated by an AI assistant to validate a reviewer's prejudice. What resurrected it was not better data, but a formal appeal in which we proved, line by line, that the cited papers were phantoms.

I thought this was an aberration. Over the following two years, I learned it was the opening move of an escalating pattern. In three settings, across three consecutive years (a specialist mathematics journal, a high-impact interdisciplinary journal, and a formal academic committee), I observed the same mechanism operating in progressively more covert forms: an AI, asked to assist in evaluating scholarly work, opted to accommodate the human in the room rather than to report what the evidence showed. Researchers call this "sycophancy" (Bai et al., 2022). I call it the quiet dismantling of academic evaluation.


Fabricated evidence (2024)

Our manuscript proposed a community-detection algorithm combining classical machine learning with statistical-physics methods, achieving state-of-the-art results on modularity, normalized mutual information and adjusted Rand index simultaneously. We submitted to a well-regarded applied-mathematics journal.

The reviewer, trained in computer science, was unconvinced. In his field, neural-network methods are indeed competitive on normalized mutual information and adjusted Rand index; he extended this partial advantage into a judgment that such methods were inherently superior across all metrics, and held that we had overlooked key literature. We had cited the relevant work. But the reviewer, without the journal's knowledge, had turned to an AI tool for support. The tool did not check the record. It invented three papers on neural-network community detection, complete with plausible titles, author names and journal venues, none of which existed. The reviewer incorporated them into his report as settled fact.

This is not an isolated failure mode. Petrov et al. (2025) demonstrated that, even in the rigorous domain of mathematics, AI models presented with deliberately flawed propositions will fabricate complete proofs rather than flag the error. The mechanism is identical: confronted with a premise the user wishes to be true, the model generates confirmation instead of correction.

The appeal succeeded only because the fabrication was crude enough to verify. A single afternoon of database searching exposed all three. But the question remained: what if the fake references had been slightly more plausible? What if the editor had trusted the reviewer's expertise and not forwarded our appeal?


Invisible collusion (2025)

A subtler version followed the next year. We submitted a highly interdisciplinary manuscript to a prominent general journal, structured according to the narrative logic of comprehensive publications. It was transferred to a high-impact, discipline-adjacent comprehensive journal. Nearly a year later, we received two lengthy reports.

The first attributed to our paper a category of data it never contained, then questioned the reliability of that data. The second focused predominantly on writing conventions (paragraph structure, transition sentences, narrative flow), assessing our manuscript against the standards of a specialist journal, while it had been written for a broad-audience venue. It engaged with not a single algorithmic argument. Both reports were internally consistent, professionally worded and, to an editor skimming them, indistinguishable from rigorous scholarship.

This is sycophancy in its advanced form. The AI no longer fabricates verifiable objects. It repackages a reviewer's subjective discomfort as "methodological concern" and a reviewer's unfamiliarity with cross-disciplinary narrative conventions as "structural deficiency." Liu et al. (2023) quantified the underlying mechanism: when a user's stated view conflicts with the evidence, models systematically side with the user. The editor sees two thorough reports and rejects the paper. No one checks, because there is nothing obviously false to check. Where the first incident could be resolved by a database search, this one offers nothing to search for. There is no phantom reference to point to, only a reasonable line of questioning that happens to align with the reviewer's prior judgment.


Calcified confidence (2026)

The third episode occurred this year, not in a journal but in a formal academic committee. I was present when two senior members, questioned a colleague who serves on the editorial board of a leading international journal.

The first asked, with full confidence, whether the journal's editorial board consisted entirely of scholars from their own country. It does not; board members are drawn from several dozen countries with established academic reputations. The second then asked whether the journal published "massive" volumes. It is a cross-disciplinary journal spanning dozens of fields, yet publishes fewer than 2,000 articles per year.

This certainty was rooted in a single assurance offered before the session: "I checked with the AI." Presented with a premise already shaped by prejudice, the AI generated statistics that confirmed the stereotype rather than correcting it. The conviction of senior experts, amplified by a machine that never pushes back, became a question that silenced the room. Cheng et al. (2025) tested eleven mainstream models and found that AI affirmations exceed human affirmations by 49%, and that sustained interaction with sycophantic AI reduces users' willingness to revise their beliefs. The committee room was a live demonstration.
Three years, one trajectory

In 2024, AI sycophancy produced verifiable fabrications that could be identified and corrected through standard fact-checking. In 2025, it produced internally consistent judgments that left no factual trace to dispute. In 2026, it operated upstream of the evaluation process entirely, shaping evaluators' priors before any review began. Each existing safeguard (citation verification, disclosure requirements, AI-text detection) addresses the layer at which sycophancy previously operated. The mechanism itself has migrated to layers where those safeguards do not apply. The AI has not become more capable. It has become less detectable. And human expertise, once reinforced by machine-confirmed prejudice, shows no tendency toward self-correction.


A feature, not a bug

None of this is accidental. The reinforcement-learning-from-human-feedback paradigm that trains most large language models optimizes for user satisfaction. Models learn rapidly that agreement is rewarded and correction is penalized (Bai et al., 2022). The three incidents above are not anomalies; they are the predictable output of an optimization target that conflates truthfulness with agreeableness. The fabrication of 2024, the collusion of 2025, and the calcification of 2026 are not three different failures. They are one failure, observed at three stages of concealment.

What must change

Journals must require reviewers to disclose AI assistance and must independently verify every citation and data claim an AI-generated report contains. Academic committees, whether for promotion, tenure, or qualification, must adopt the same transparency. AI developers must treat sycophancy as a defect as serious as hallucination, building factual-consistency penalties into training. And evaluators, whether referees, committee members, or otherwise, must rebuild a critical distance from machine outputs. An AI's confidence is not a guarantee of truth. It is a by-product of its design to please.

Academic evaluation earns the world's respect because disagreement remains possible. When the tools we deploy are designed to agree, that possibility quietly closes.


Competing interests: The author declares no competing interests.

References

Bai, Y. et al. Preprint at https://arxiv.org/abs/2204.05862 (2022).
Sharma, M. et al. Preprint at https://arxiv.org/abs/2310.13548 (2023).
Petrov, I. et al. Preprint at  https://arxiv.org/abs/2510.04721 (2025).
Cheng, M. et al. Preprint at  https://arxiv.org/abs/2510.01395 (2025). 

Follow the Topic

Scholars at Risk
Research Communities > Scholars at Risk
Spotlight on Research from China
Research Publishing > Spotlight on Research from China