Why "100% AI-Generated" Is a Verdict, Not a Fact

When AI is trained on the achievements of all humanity, and used polish expression, on what grounds is the entire work labelled "a product of AI"? Where exactly should this boundary be drawn? This article proposes a three-element framework of "contribution — control — accountability".
Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Introduction

The academic world in 2026 is staging a structural absurdity: a doctoral candidate who spent four years on experimental design, mathematical derivation and code reproduction had their thesis flagged by a detection system as showing "artificial-intelligence involvement with extremely high confidence" — simply because, at the final-draft stage, they had entrusted a large language model with translating their native-language draft into academic English.

The technology behind this controversy is not mysterious. To comply with the transparency requirements of Article 50 of the EU AI Act, vendors such as Anthropic have begun embedding statistical invisible watermarks, based on probability-distribution perturbation, into model outputs: as the model generates text word by word, it applies a slight, key-traceable probabilistic tilt to candidate tokens, so that the whole text carries a traceable "algorithmic fingerprint" at the statistical level. The crux is this — the mechanism cannot distinguish between "a translation of human thought" and "machine-authored ghostwriting". So long as every English token has passed through the model's probabilistic sampling, the translation is, by mathematical definition, indistinguishable from fully automated generation.

What should have been a technical matter internal to editorial offices has thus abruptly escalated into a theoretical question: when AI is trained on the language and intellectual achievements of all humanity, and a human being merely borrows it to polish expression, on what grounds is the entire work labelled "a product of AI"? Where exactly should this boundary be drawn, and who is to draw it? For editors of academic journals, this is no longer a distant topic of ethical conversation, but a real adjudication that may land on the triage desk on any given day.


1. Three theoretical dislocations in the boundary dispute

(1) Ontological dislocation: mistaking the form of expression for the substance of thought

The logic by which technology vendors defend watermarking runs as follows: "every word of the translation was chosen by the model, therefore the text is a machine artefact." The blind spot of this claim is that it conflates the ontological distinction between language as "vessel" and thought as "content". The history of human technological evolution is, at bottom, a history of inventing tools to reduce the friction of expression: from the writing brush to the typewriter, from spell-checking to grammar correction — we have never required a pen manufacturer to embed hidden marks in the handwriting, declaring that "this line was not written entirely by a human hand". The reason is self-evident: what gives a paper its soul is the researcher's years of contemplating the unknown, the rigorous selection of data, and the intellectual sparks struck in logical impasses — not the mechanical act of token selection.

Equating "form of expression" with "substance of thought" is, philosophically, a category error. A scholar's core hypotheses, experimental design and interpretation of results do not change their epistemic status merely because the grammatical shell of the sentences has been machine-polished. Polishing alters the "surface structure" of the text, not the "deep structure" of the knowledge.

(2) Epistemological dislocation: mistaking statistical signals for causal evidence

What AI-detection tools (including watermark decoding and perplexity analysis) output are, in essence, probabilistic signals, not findings of fact. Comparative studies show that the accuracy of mainstream detectors fluctuates roughly between 25% and 60%, with false positives and false negatives coexisting; empirical research from institutions such as Stanford University has further revealed systematic bias against non-native English-speaking authors. This is because well-trained academic writers produce regular sentence patterns, restrained diction and naturally low textual perplexity — features that overlap heavily with the output characteristics of large models.

This means that a detection score is, in legal and ethical terms, at most a "lead", and never constitutes "evidence". Elevating statistical correlation directly into causal determination is the most elementary error in scientific methodology (see my previous comment, "Correlation Is Not Causation, But Our Review Process Pretends Otherwise"); and when a journal institutionalises this error — for example, by using a "detection rate" as the trigger for rejection or investigation — it is using an epistemologically unreliable instrument to carry out academic adjudications of the gravest consequence. The core commodity of scholarly publishing has never been words themselves, but accountable knowledge claims; the adjudication of knowledge claims should follow a standard of evidence higher than word-frequency statistics.

(3) Ethical dislocation: the broken accountability chain and asymmetric "digital enclosure"

The most ironic aspect of the boundary dispute is the asymmetry of the accountability structure. Large models are trained on the entire corpus of the internet — a corpus that includes countless scholars' papers, translations and intellectual achievements — yet AI systems, or the technology service providers behind them, have never fulfilled any obligation of "attribution" to these knowledge contributors. And when the same scholar borrows the model to polish their own words, the output is stamped with the hidden mark of "machine artefact", and ownership of originality is unilaterally stripped away. Knowledge flows into the model without boundaries, yet flows out behind high fences; politically, this asymmetry constitutes a new form of "digital enclosure": platforms, in the name of compliance, shift the full costs of regulation and of false accusations onto the downstream individuals with the least bargaining power — above all, researchers who are not native English speakers.

Meanwhile, the accountability chain faces rupture at the other end: if AI is allowed to generate research content substantively and without constraint, then when a paper contains fabricated data or erroneous conclusions, no one will be able to bear academic responsibility. An AI system cannot meet the fundamental requirements of authorship — it cannot answer for the work, cannot be called to account, and cannot suffer the reputational consequences of retraction (CNKI, China's largest and best-known publishing platform, has issued a "Statement on the Handling of Papers in which Artificial Intelligence (AI) Is Listed as an Author" ). This is precisely the ethical ground on which the boundary must exist: the boundary is not meant to exclude technology, but to safeguard the possibility of accountability.


2. Redrawing the boundary: from "textual provenance" to "accountable intellectual contribution"

Since drawing the line by "who generated the text word by word" has collapsed in both theory and practice, society needs a new principle of demarcation. This article proposes a three-element framework of "contribution — control — accountability":

First, the source of intellectual contribution. The key question is not "who wrote the sentences" but "who introduced new intellectual content". Research questions, hypotheses, methodological design, data interpretation, argumentative structure and conclusions — these constitute the substantive core of scholarly contribution. If these cores originate with the researchers themselves, then whatever tool polishes the textual shell does not alter the attribution of the work; conversely, if the core argument is generated by a model and presented under a human name, then however "human" the prose may be, the boundary has already been crossed.

Second, the substantiveness of human control. Is the tool's intervention under substantive human control — can the user understand, verify, modify and overrule the tool's output? In the polishing scenario, the author checks the accuracy of the translation sentence by sentence and retains final authority over the text; the chain of control is intact. In the "one-click literature review" scenario, the user is often unable to verify the authenticity of every citation; control has already slipped away.

Third, the traceability of accountability. Is there an identifiable human subject who can answer, to the very end, for every knowledge claim in the text? Accountability cannot be transferred to machines; this is the cornerstone of trust on which the scholarly community operates.

Viewed against these three elements, the judgement that "polishing equals AI generation" collapses of its own accord: polishing introduces no new intellectual content, the author retains complete control, and the accountable subject is clear. Normatively, it is no different in kind from spell-checking software — what differs is only the efficiency of the tool, not its ethical rank.


3. How should editors think?

It is precisely along these lines that Dr @Gino D'Oca , Editor-in-Chief of Humanities & Social Sciences Communications, recently issued to the editorial community a risk-based framework for assessing AI use, sorting the myriad use cases into a "green — amber — red" spectrum.

Green zone (permitted): assistive, low-risk AI use. The criteria are: AI supports expression, organisation or efficiency without influencing scientific, scholarly or evaluative judgement; the use is reversible and verifiable; it introduces no new intellectual content; and accountability remains clearly and continuously human. Language polishing, translation, suggestions on manuscript structure, data cleaning and deduplication all fall into this category. This provision in effect formally declares, at the policy level: polishing is not ghostwriting.

Amber zone (exercise caution): evaluative or interpretive AI use. When AI begins to intervene in reasoning and critique — suggesting analytical approaches, drafting explanatory summaries, identifying patterns in exploratory data, recommending statistical tests — it introduces new intellectual content, which must be verified and placed under human oversight, with human judgement and accountability "demonstrably" present. Such use is permitted, but on the premise of transparent disclosure.

Red zone (not permitted): AI replacing scholarly judgement, or use lacking transparency. Generating hypotheses, analyses or conclusions and passing them off as human achievements; fabricating data and citations; delegating peer review to large models; ceding authorship and accountability to AI systems — these uses are opaque, their outputs unverifiable, their accountable subjects hollowed out; none of them is permitted.

The deeper wisdom of this framework is that it wrests the boundary question out of the hands of the "statistical fingerprint of the text" and returns it to the "accountability of human judgement". It presupposes a key position: what detection tools measure is not "contribution" but mere "exposure"; and what editors must adjudicate is never which tools a text has touched, but whether the intellectual contribution within it can be traced to an accountable human subject. At the same time, this editorial policy is underpinned by four core expectations: human accountability cannot be transferred to AI systems; AI may support but must not replace scholarly judgement; transparency about AI use helps to build trust; and the use of AI tools must uphold confidentiality and data protection.


4. Journal practice

Theory must ultimately land on the editorial workflow. In light of the framework above, this article offers the following recommendations for journals in the two directions of AI application and AI detection:

First, detection signals must not serve as sole evidence. Any AI-detection report — whether from watermark decoding or perplexity analysis — may only initiate an "enquiry procedure", never directly a "penalty procedure". Editors are obliged to give authors the opportunity to explain and to provide evidence: research notes, raw experimental data, writing-iteration records and version histories are all far more reliable evidence of provenance than word-frequency statistics. This is the minimum that procedural justice requires of academic adjudication.

Second, apply "bias correction" for non-native-speaking authors. Given that detection tools have repeatedly been shown to misjudge non-native English writers systematically, journals should maintain institutional caution towards high detection scores from this group during initial screening, lest a language disadvantage evolve, in the algorithmic age, into a new form of publication discrimination. AI polishing is, in origin, an inclusive "instrument of linguistic equity"; journals should not conspire with misjudgement mechanisms to turn it back into an "original sin" borne by minority groups.

Third, disclosure over prohibition; gradation over blanket rules. Journals should establish clear guidelines for disclosing AI use, encouraging authors to state voluntarily the tools, purposes and verification methods involved, and should handle cases separately according to the green — amber — red logic: green uses require only a statement, with no obstacles; amber uses require an account of the verification process; red uses are inadmissible whether disclosed or not. A blanket ban would only breed more covert circumvention techniques, driving the grey zone underground.

Fourth, the boundaries on the side of editors and reviewers are equally rigid. Reviewers must not upload unpublished manuscripts to any generative AI tool — this is the bottom line of confidentiality obligations and data protection; if any part of a review report has been supported by AI tools, this should be transparently declared in the report. AI may be used in administrative tasks such as plagiarism screening, format checks and initial conflict-of-interest screening, so as to relieve the editorial workload, but review judgements and acceptance decisions must remain in human hands.

Fifth, beware of "defensive bureaucracy". Under the risk-averse mentality of "better to kill by mistake than to let one through", journals are prone to using algorithmic verdicts as a shield from liability. But the dignity of the editorial profession lies precisely in judgement; outsourcing judgement to a detector is, epistemologically, the same abdication as outsourcing writing to a generator.


Conclusion

The dispute over AI's boundary is a technical question on the surface, a philosophical question at depth, and ultimately an institutional question. When we ask "to whom does a polished text really belong", what we are really asking is this: in an age when machines can speak fluently in every language, what is scholarly publishing still exchanging? The answer is not words, but knowledge claims made by named human subjects who are willing to bear full responsibility for them. Words can be outsourced; responsibility cannot.

The true boundary, therefore, lies neither in the statistical fingerprint of a watermark nor in the confidence interval of a detector, but in the accountability of human judgement. For editors, the way to hold this boundary is not sharper detection tools, but more mature frameworks of judgement, more proper investigative procedures, and fidelity to the simple order that tools serve thought. Polishing is not ghostwriting; detection is not adjudication; and judgement — forever — belongs to humans.

Please sign in or register for FREE

If you are a registered user on Research Communities by Springer Nature, please sign in

Follow the Topic

Moral Philosophy and Applied Ethics
Humanities and Social Sciences > Philosophy > Moral Philosophy and Applied Ethics
Artificial Intelligence
Mathematics and Computing > Computer Science > Artificial Intelligence
Philosophy of Artificial Intelligence
Humanities and Social Sciences > Philosophy > Philosophy of Science > Philosophy of Technology > Philosophy of Artificial Intelligence