Life in Research

What rebuilding a systematic review taught me about using AI responsibly

A missing audit trail forced me to rebuild a seemingly finished meta-analysis. The experience changed how I use AI and which scientific decisions I keep under human control.

While rebuilding a review of biomaterial-based treatments for erectile dysfunction after nerve injury, I found an old conference abstract in our files. It contained precise study counts, pooled effects and percentages describing possible mechanisms. My first reaction was relief: perhaps a difficult part of the work had already been done.

Then I looked for the extraction sheet, the record of how the studies had been selected, the analysis code and the references behind each number. I could not find a trail that I could verify. The figures might have been correct, but I had no defensible way to know.

We marked them as unverified and started again. Rebuilding the work cost time and briefly felt like moving backwards. It also taught me more than a smooth analysis would have done. As a doctoral medical researcher in urology, I now use AI to help organize records, test code, compare documents and shape drafts. Its greatest risk in my work is not always an obvious error. It is premature coherence: an unfinished process can suddenly look complete.

Five decisions have therefore become my checkpoints.

Is the review worth doing?

AI can map previous reviews, suggest search terms and turn a broad interest into a polished protocol very quickly. It cannot decide whether a new review will add knowledge that matters.

Before expanding a project, I ask whether recent reviews already answer the question, whether enough independent groups of patients or experiments exist, and whether the outcomes address a real clinical uncertainty. I also ask whether the project would still be useful if the studies proved too different to combine. A blank patch on a literature map is not automatically an important evidence gap. Sometimes the responsible decision is to narrow the question, change the design or stop.

Am I counting studies or papers?

In a single-cell cancer project, my team found that a paper that appeared to provide independent validation reused an earlier public dataset. Counting both papers separately would have made the finding look more reproducible than it was.

This is why the Cochrane Handbook treats the study, rather than each report, as the unit of interest. AI can match author names, repository identifiers and sample descriptions, but a dataset's family tree may be hidden in a supplement or a single sentence in the methods. I keep a paper-cohort-dataset ledger and attach a source to every judgement about independence. When the lineage remains uncertain, I record possible overlap instead of converting uncertainty into another study.

What does this number mean clinically?

A polished forest plot, the familiar chart used to display results in a meta-analysis, can make incompatible outcomes look coherent. In reviews of sexual-function treatments, for example, improvement measured while patients are receiving treatment is not the same as spontaneous recovery after treatment stops. Measurements taken at different follow-up times may also answer different clinical questions.

Before extracting a number, I write its meaning as a sentence: in which patients, compared with what, for which outcome, at what time, and measuring what kind of effect? If two studies produce different sentences, I do not let a shared spreadsheet heading erase the difference. Statistical software may accept the inputs, but the clinically honest result may be a structured explanation rather than one pooled estimate.

Is the search actually complete?

In one recent review, searches of PubMed, Web of Science and Scopus produced 9,643 records. After exact duplicates were consolidated, 5,032 candidate groups remained. The spreadsheet looked substantial. Yet another database, an independent review of the search strategy and final duplicate checks were still unresolved, so we did not begin formal screening.

That pause also felt like moving backwards. In reality, it protected the review from a false start. The PRISMA 2020 statement makes the search and selection process visible because readers need to understand how the evidence set was formed. I now retain every native query, reconcile the result count shown by a database with the exported file, record access limitations and send uncertain duplicates for human review. "Not yet searched" is more informative than a silently empty cell.

How far can the evidence carry the claim?

The final decision concerns language. An absence of reported harm is not proof of safety. An association is not causation. Reusing a dataset is not external validation. Large differences between studies do not disappear because a different statistical model is selected.

For each major finding, I write three short statements: what the evidence shows, the strongest claim it supports, and the claim it does not support. AI can help test the wording inside that boundary, but I check every number against its source. This small routine has repeatedly stopped confident prose from outrunning uncertain evidence.

Progress looks different now

These checkpoints have slowed, narrowed and redirected some of my projects. In the single-cell project, I postponed installing a full analysis environment until the question and usable datasets were clear. In the biomaterials review, we rebuilt the analysis instead of recycling attractive numbers. Both choices cost visible progress at first, but probably prevented much more wasted work.

I no longer define efficiency as the shortest time to a finished-looking draft. A useful AI-assisted workflow should leave receipts: the query that was run, the file that was exported, the source behind a judgement and the boundary around a claim. The Research Communities guidance on AI places responsibility for accuracy and integrity with the human author. That is the kind of partnership I want with AI: fast enough to help, transparent enough to audit, and never a substitute for deciding what the evidence deserves.

AI-use disclosure

OpenAI Codex assisted with organizing verified project notes, structuring the article and editing draft language. I checked each example and link and take responsibility for the final text.

Competing interests

The author declares no competing interests.

Author bio

Xiancheng Du is a medical researcher and doctoral candidate in Clinical Medicine (Urology) at the Department of Urology, Zhongda Hospital, School of Medicine, Southeast University, Nanjing, China. His work spans evidence synthesis, regenerative medicine and computational urology.

Header photograph: Rui Hou.