On the Rapid Integration of AI Tools into the Peer Review Process
Executive Summary
AI tools may help address the growing demands of scholarly peer review by checking citations, evaluating conclusions, and improving clarity, but their usefulness depends on the task and the quality of human oversight. Journals should establish explicit guidelines governing AI use by both authors and reviewers, including requirements for transparency, appropriate prompting, and evaluation of AI-generated feedback. Although AI may outperform humans on some routine technical checks, it remains less reliable across disciplines and should not independently determine a paper’s novelty, significance, or publication outcome. Effective integration therefore requires regulations developed by both AI specialists and experts in scholarly publishing, with humans retaining responsibility for intermediate and final decisions.
Introduction
Will AI help or hurt peer review as traditionally practiced by scholarly scientific publications? This question is addressed in a recent AAAS Science article, “Ready for Robot Reviewers?”
The answer depends to a great extent on the part of the peer review process to which AI tools are applied:
Checking citations and references.
Assessing whether a paper’s conclusions hold up.
Ensuring that a paper’s language is clear and grammatical.
These applications require different skills and capabilities.
Guidelines for AI-Assisted Review
One challenge is that, as the number of papers submitted for publication continues to rise, so does the pressure on editors to find enough qualified reviewers. The Science article describes how AI tools are making inroads into the process. But just as I would not criticize researchers who use AI tools to analyze data and draw preliminary conclusions, I would not criticize those who use AI tools in the review process—as long as that process is overseen and managed by humans exercising professional judgment.
If we are going to allow AI into the review process, editors need to be specific about its “dos” and “don’ts,” keeping in mind that (a) reviewers are presumably selected for their expertise, (b) reviewers agree to use their expertise to oversee and evaluate what AI tools suggest, and (c) reviewers know how to use AI tools effectively.
That last item is critical. In my own use of AI tools such as ChatGPT to help edit my writing, I have found that creating prompts to guide an AI application is a skill that evolves with practice, especially when one uses general-purpose AI tools in the editing process.
The more structured and rule-based the editing process becomes—and the more detailed and explicit the initiating prompt—the more useful the resulting edit is likely to be.
This suggests that editors seeking volunteer reviewers for research articles must be explicit about what is and is not allowed in an AI-assisted review. It also means that a journal’s editorial guidelines must include explicit directives governing how AI may be used by both authors preparing articles and reviewers evaluating them.
Evaluating Novelty and Significance
Clearly, these two uses need to be coordinated. But how does one use AI to assess “novelty or significance,” especially when so many articles are multidisciplinary and require knowledge of several fields? Of course, one can ask the same question about selecting reviewers for multidisciplinary papers. This is one reason editors might seek multiple reviewers for an individual paper.
Eventually, we may need to face a reality: situations will arise in which AI tools—especially those capable of researching and evaluating work across many different fields—are more technically knowledgeable than the reviewers themselves.
This does not necessarily mean that AI tools are more qualified to make judgments about “novelty or significance.” It does mean that reviewers using AI tools should require those tools to explain the specific steps taken in the review process in a way that is clear and understandable. Only then can a reviewer make an informed judgment about any preliminary, intermediate, or final feedback produced by the AI.
Strengths and Limitations
One mistaken assumption would be that human reviewers are always preferable to AI reviewers. Humans are fallible, and some human reviewers concentrate on minor or ephemeral issues while skimming. The Science article emphasizes that AI tools can “shine” when performing the routine, tedious checks for correctness that time-pressed human reviewers often skip—something I have also found in my own writing.
The Science article reports positive results from one experiment in which computer science conference attendees used Google’s Gemini-based “Paper Assistant Tool” to improve drafts before submission. It also describes the bioRxiv preprint server, where authors can obtain technical checks before submitting their work.
At the same time, several studies discussed in the Science article have shown that LLMs can be more error-prone when dealing with practices or disciplines outside a particular paper’s focus. Other researchers question the tendency of some AI tools to “nitpick” less important technical details while ignoring more important findings. When it comes to reviewing, then, neither humans nor AI tools are perfect. But we knew that already, right?
Conclusion
Given the rapid advances in both journal-publishing practices and relevant AI tools, it is clear that we have much to learn about how best to incorporate AI into the broader communication processes of the scientific research lifecycle. Still, at least three things seem clear to me about this rapidly evolving area:
Regulations governing the use of AI tools should be developed by people who understand how those tools operate.
Those regulations should also involve people who understand the processes to which the AI tools will be applied.
Humans must remain involved in intermediate and final decision-making concerning those processes.
Copyright © 2026 by Dennis D. McDonald. The Executive Summary was created by ChatGPT as was the graphic based on my request for a green eyeshade wearing editor skeptically reviewing a submitted manuscript.


