Where AI could help medical journal peer review
My view is that there is a useful role for AI medical journal peer review. But the evidence is much stronger for some individual tasks than it is for improving the journal’s overall performance.

AI in medical journal peer review is attracting a lot of attention. Journals are receiving more manuscripts, editors are struggling to find reviewers, and there are plenty of tools promising to make the process faster.
I wanted to look at how much evidence sits behind those promises, and where AI might be useful in practice.
I’ve put together two papers. AI in medical journal peer review: what the evidence shows reviews the research, commercial tools and policies. Assist, not assess builds on that review and proposes a framework for using AI in the work around the reviewer.
My view is that there is a useful role for AI here. But the evidence is much stronger for some individual tasks than it is for improving the journal’s overall performance.
What has actually been measured
The first paper separates measured results from proposals and commercial claims.
We found no published study showing that AI has shortened turnaround time or cleared a backlog at a medical journal. That does not establish that AI cannot help. It means the outcome publishers are most interested in has not yet been demonstrated in the literature covered by the review.
There are results worth paying attention to.
In one study, an LLM compared reported trial outcomes with registry records in about two minutes, compared with 27 minutes by hand. At The BMJ, models extracted CONSORT reporting items from 50 trial manuscripts in 38 to 75 seconds, with accuracy reaching 93.4%.
Image screening has also identified problems before publication. The American Society for Microbiology found duplications in 3.9% of 2,627 accepted manuscripts screened over one year.
These are useful findings. But they measure the speed or accuracy of a task. They do not tell us whether an editor spent less time on the manuscript, whether fewer reviewer invitations were needed, or whether the author received a decision sooner.
A tool might save time on one check while creating extra work elsewhere. Someone still needs to interpret its flags and deal with mistakes.
Why I would keep assessment with people
The evidence for using AI as the reviewer is less encouraging.
Across 300 manuscripts at three ophthalmology journals, human reviewers recommended rejection for 73.33%. Each LLM tested recommended rejection for 2.00%. Other studies in oral surgery and radiology also found that models were reluctant to reject manuscripts.
There is also a problem with manipulation. In one experiment involving medical manuscripts, hidden instructions changed the recommendation in 84.4% of 720 simulated reviews.
These findings concern the models and conditions tested. They may change as models improve. For now, though, they give us good reasons to keep scientific assessment and editorial decisions with reviewers and editors.
Fluent review text is not enough to establish that a model can judge a study reliably.
The work around the review
The second paper looks at the surrounding workflow.
Finding reviewers is an obvious difficulty. Across 21 BMJ Group journals, only 35.2% of 257,025 invited reviewers agreed to review. That leaves editors sending further invitations and waiting for responses.
There are also submission checks, integrity screening, reminders and preparation of material for editorial decisions.
We do not have a measured breakdown showing exactly how much delay each step contributes at a medical journal. The argument that logistics and preparation account for a substantial share is an inference. I think it is a reasonable place to investigate, but it still needs testing.
The framework covers seven steps:
1. Checking that a submission is complete.
2. Screening for integrity problems.
3. Preparing a brief for the handling editor.
4. Finding suitable reviewers.
5. Supporting reviewers with check results and feedback on their own drafts.
6. Managing invitations, deadlines and reminders.
7. Bringing reviewer reports together for the editor.
The evidence varies across these steps. Integrity screening has practical results behind it. Reviewer matching tools are widely available, but we found no medical journal outcome study showing that they secure reviewers faster. Process management and decision preparation remain proposals in this framework.
Keeping those distinctions visible matters. A sensible use case is not automatically a proven one.
How I would decide what to try
Before introducing a tool, I would want to know whether a person can check its output, who makes the final decision, and what happens when it is wrong.
A flag for a duplicated image gives someone a specific thing to inspect. A missing declaration can be checked against the submission. An opinion about a paper’s novelty is harder to verify.
The journal also needs a secure environment for manuscript content, a way to identify manipulation attempts, and clear disclosure of where AI is used. Editors and reviewers need to understand the tools well enough to use them and question their output.
Then we need to measure the result.
The framework proposes recording a baseline and introducing one step at a time, or staggering introduction across journals or sections. Measures should include time to first decision, time to secure reviewers, invitations per completed review, and editor and staff minutes per manuscript.
False flags, review quality and author and reviewer satisfaction also need monitoring. A faster process would be a poor result if it created more errors or worse reviews.
I think there is enough evidence to justify testing specific forms of AI assistance carefully. There is not yet enough to promise that they will solve the peer review backlog.
The two papers explain the evidence and the proposed approach in more detail, including references and limitations. I’d welcome feedback from editors, reviewers and publishing teams, particularly from anyone already measuring what changes when these tools are introduced.

