This is the third post in a series on the new questions raised by AI. My previous posts discussed AI risks and the moral status of AI agents. In this post, I look at how AI is set to revolutionise the peer review process at the heart of modern science.
In 1936, Einstein and Nathan Rosen submitted a paper to Physical Review titled “Do Gravitational Waves Exist?” The editor sent it to an anonymous referee who criticised it. Einstein reacted angrily:
We … had sent you our manuscript for publication and had not authorized you to show it to specialists before it is printed.1
He refused to answer the referee’s criticisms and withdrew the paper. This episode is very foreign to the modern experience of the peer review process. It shows how relatively recent the norm of peer review is in science. External peer review only became general after the Second World War, taking a form close to its modern use over the following decades.2
Today, only scientific contributions that have been “peer reviewed” are typically considered proper scientific publications. They have been vetted by academics in the relevant fields, typically two or three, sometimes many more for leading outlets. Based on the feedback from these reviewers, the editor processing the paper decides either to reject the submission (because its approach is seen as flawed, or correct but with too minor an intellectual contribution), or to invite a revision in which the authors try to address the comments and concerns of the reviewers and get them to give the green light to the editor for the paper to be finally accepted.
The arguments in favour are obvious: quality control by equally qualified peers is bound to improve the quality of the outputs. This process is certainly a major factor in the success of science in producing much better arguments and theories than other fields of argumentation where such stringent quality control does not exist (for instance, political discussions). But peer review also has its critics; the peer review process is not perfect and features well-known biases.3
In a matter of a few months, the technical capacities of frontier AI models have upended the conditions that made the prevailing conventions in research stable. Here I look at how AI is likely to impact peer review. My view is optimistic. I think it is going to change it for the better.
The benefits of AI in peer review
Many people have raised concerns about AI peer review. Some of these concerns are justified by the experience of earlier models, which produced generic reviews of middling quality.
Recent models, such as ChatGPT 5.6 and Fable, are, however, a different beast. It is important to appreciate what they can do to understand how they will revolutionise science.
Benefit 1: depth and breadth of the review
The conceptual work that frontier models can now do is quite stunning. In a few minutes, they can read a document, analyse it in depth and identify small mistakes that would take a human much longer to find.
The depth of their analysis is just mind-boggling. For theoretical papers, they can check mathematical proofs in a fraction of the time the same work would take a human. For empirical papers, when the relevant data and code are available, they can inspect calculations and identify discrepancies that would require incredible attention to detail for a human to find.4
Their breadth of knowledge is equally remarkable. Frontier models have access to a range of knowledge far beyond that of any individual human. They can systematically check references and factual assertions in a text, assess whether they are accurate and flag noteworthy work or facts that may have been left out.
If you combine these two aspects and ask a frontier model to write a report about a paper, it can produce something incredibly insightful.
Benefit 2: faster feedback
The lightning speed at which AI feedback can be provided is, in itself, a game changer. Assessing research outputs carefully can be very hard. Scientific studies are often technical, mathematical proofs take time to understand, and statistical studies can have a lot of moving pieces that need to be checked carefully.
For that reason, peer review can greatly lengthen the production of scientific outputs. Some disciplines, like computer science, tend to be fast, as competition is rife on very similar questions and techniques. But in other disciplines, it can take a very long time. In economics, it can easily take six or seven years between the start of a project and its publication in a leading journal.5
AI capabilities greatly reduce this problem by making the hard work of checking the technical aspects of the paper fast.
Benefit 3: less bias
Science is a social activity, and peer review is a social judgement. It comes with all the problems of social judgements, not always far from our experience of judgement by our peers in high school. Reviewers are humans with their preferences, prejudices and interests. A very good paper might face the unfair ire of a reviewer because the reviewer does not like the result, does not like the author, or has a competing theory to defend.
In science, peers are your colleagues and your competitors. They are typically those who published in the same area before. They may have an interest in defending their turf and their reputation. Papers, as brilliant as they may be, that state “Previous research in the area is bunk, here is a much better approach” are likely to go down with reviewers like a lead balloon. It is a mistake experienced by many brilliant but overconfident graduate students who think they will revolutionise a research field with their great ideas, only to see their paper rejected until they water down their claim substantially.
Max Planck famously quipped that new scientific truths often triumph because their opponents eventually die and a new generation grows up familiar with them.6 A recent study of the premature deaths of eminent life scientists provided support for Planck’s quip. It found that scientists who were not collaborators of the star scientist published more afterwards. This increase seemed disproportionately driven by outsiders, scientists bringing a different scientific corpus into the field. This is consistent with the idea that leading scholars may actually slow the progress of competing approaches.
Frontier models have no career, reputation or theory at stake in the research they review. They can of course reproduce biases embedded in their training data, but they do not have idiosyncratic preferences, personal grudges, friendships or career interests tied to specific authors. In that sense, they can be much more neutral judges.
Benefit 4: less statistical discrimination
Peer review is also imperfect because peers have limited time (typically a few hours) to understand something that the authors might have been working on for several years. One consequence of this asymmetry of information between authors and reviewers is that reviewers (and editors) will rationally use any available signal to form a view about the likely quality of the output they are trying to assess. What signals? The main ones are the known record of the author and his or her institution. The (rational) use of these pieces of information can lead to two types of well-known biases: prestige bias and familiarity bias.
Prestige bias: researchers with prestigious track records or positions at prestigious institutions are more likely to see their output viewed in a positive light in the peer-review process. This is certainly one of the most pervasive forms of discrimination in academia (though perhaps unsurprisingly, leading associations and organisations filled with academics at the top of the prestige ladder rarely discuss it).7

Suppose a reviewer is not sure they understand a paper. If that paper is written by a Nobel Prize winner, the reviewer is more likely to think that their lack of comprehension is due to their own limitations (“he must know something that I don’t”). If that same paper is written by a grad student, the reviewer is more likely to think that the problem lies with the paper. This mechanism is one possible explanation for the Matthew effect, proposed by the sociologist Robert Merton: scientific recognition tends to accrue disproportionately to those who already have it.8
In a recent experiment, economists sent the same finance paper to academics under three conditions: with Nobel laureate Vernon Smith shown as the author, with a junior and still little-known researcher shown as the author, or with no author information. The result was striking. A large majority of reviewers rejected the paper when the little-known author was shown, while the same paper received much more favourable recommendations when Smith’s name was shown.9
In another experiment, reviewers at an academic conference in computational neuroscience were either shown or not shown the names of the authors of a submission. Non-students and submitters from top institutions received higher scores when reviewers were not blinded, and blinding substantially reduced these gaps (without worsening the ability of scores to predict later citations and publication outcomes).10
Interestingly, one study also looked at the existence of prestige bias in LLMs. Using GPT-4o-mini, the researchers varied the information given about the authors and found that authors from prestigious institutions were favoured. The level of bias appeared, however, lower than in studies of human reviewers. While there were some significant effects on the scores given by the AI reviewer to papers, the effects on acceptance/rejection decisions were smaller and less robust.11 Since such bias is, in a sense, rational when the reviewer has limited ability to assess the paper itself, I would expect it to be even smaller with present frontier models, which are substantially better than the model used in that study.
Familiarity bias: Because only a minority of academic submissions can be accepted in leading conferences or leading journals, reviewers and editors do not simply have to identify whether a contribution is good; they have to identify whether the contribution is among the best being submitted. In statistical terms, they need to assess whether it is in the upper tail of the distribution. Here another rational discrimination can take place: screening discrimination.12 It is easier to identify whether something is very good if you know how to read and understand the associated signals of quality. You might, for instance, prefer to hire a graduate from your university at your firm simply because you can assess that applicant’s credentials better than if he or she had a degree from a university that you do not know. In that latter case, because you are less able to read the signal, it is harder for you to confidently place that applicant at the top of the pile.
This discrimination in academia will benefit those who are in close networks and know each other, and are therefore better able to identify which authors are known for producing good work. It can also penalise papers that are too novel in their topics or methods and for which reviewers are less certain about where they stand.
Benefit 5: less randomness
Something experienced by all editors is that reviewers’ judgements are only poorly correlated. It is very frequent for reviewers to make different recommendations, sometimes radically different ones. Studies of peer review have confirmed this fact. In one study of papers submitted to a medical journal, it was found that reviewers’ accept/revise recommendations conflicted in about 45% of cases.13 Note that if two reviewers were to throw a coin in the air to make their recommendation, they would disagree 50% of the time!14 The use of several reviewers and the editor’s overall judgement might mitigate this randomness, but variation in reviewers’ recommendations has to influence final decisions. One study looked at a machine learning conference where submissions were handled by two independent program committees and found that accept/reject final decisions differed for 25% of the cases.15
My experience as an editor aligns with these types of numbers, with reviewers sometimes having fiercely opposed views.
Peer review is therefore in part random: the outcome might depend on the luck of the draw and whether you get reviewers who happen to like your work or not. In principle, this also gives an editor some discretion to invite reviewers who are more or less likely to be sympathetic to a paper if the editor already has a preference about the final outcome.
Another consequence of the randomness of reviewers’ takes is that the reviewers’ requests to address some concern or improve the manuscript often appear somewhat idiosyncratic. Each new submission often leads to requests that seem unrelated to those of past reviewers. Scientists can also often be heard saying that they feel the revision process added unnecessary elements, simply to please the views of some reviewers.
Frontier models have the advantage of showing less randomness from one request to another, and even from one model to another. In effect, they act as synthetic reviewers, aggregating the insights of a large number of possible reviewers. Getting an AI review therefore reduces the role of chance in the editorial decision.16
The AI review revolution
Given all these benefits, there will be an AI revolution in peer review just as there will be one in scientific writing. It would not make sense not to use these tools to assess papers given their advantages and the ways in which they improve on pure human review.
Where humans are still better
What place is there for humans then? Should we simply delegate editorial decisions to AI models? Even without defending our place as humans in the system just for the sake of keeping our academic positions, there are reasons, at least at the moment, not to outsource all decisions to AI.
First, one aspect where these models are weakest is in assessing how “interesting” a result is going to be for the possible readership of the journal. This is perhaps the most difficult task for which the expertise of the editor matters: assessing whether a paper is interesting and worth presenting to the scarce attention of readers requires an understanding of what readers are interested in, what they know, what the paper adds to what they know, and how this addition will be useful to them given their goals and interests. This involves a kind of “collective mind reading”, the ability to “read the room”. It is not surprising that it is one of the hardest tasks for models to perform.
The latest models are so good that the human advantage might be slim here, but nonetheless I would trust humans to keep making these judgements.
A second possible human advantage concerns novelty. By design, an AI model may tend to judge work against patterns represented in its training data and therefore be conservative towards genuinely unfamiliar ideas. Humans are often conservative too, but a specialist with deep knowledge of a narrow area may see the interest of an original idea with greater acuity than an AI model.17
A recent study seems to support such a possibility. It found that the proposals rated as more novel by human reviewers were disproportionately funded by humans but not by GPT-4o.18 This is also a reason to keep human judgement in the review process. If AI models have a conservative bias towards unfamiliar ideas, their greater consistency than human reviewers could mean that a valuable and original paper is systematically downgraded by AI reviews.
AI-enhanced reviews
The first step in the revolution is to use AI to help reviewers assess papers: you combine the AI’s general take with the reviewer’s expertise. This is already happening informally. A recent study estimated that between 6.5% and 16.9% of review text at four major AI conferences was substantially modified or produced by LLMs.19
A reviewer can use the AI model not simply to take its view for granted, but to interrogate the paper and form a better-informed personal judgement. A large randomised experiment at a machine learning conference in 2025 found that 27% of human reviewers who received AI-generated feedback on their draft reviews revised them. A blinded evaluation then found that the revised reports were more informative than the original ones.20
AI review
If everybody uses AI, however, it makes sense for journal editors to use it in a transparent way and announce to submitters that they will be using an AI model to assess their paper. This is now starting to happen. In August 2026, the American Economic Association announced a partnership with an AI review tool, Refine, to use its AI-assisted technical verification tool on papers that have received, or are under serious consideration for, a revise-and-resubmit decision. The AEA describes this AI review as a technical check only: it will not be used to assess the contribution or importance of the paper. Authors will also receive the AI-generated comments to help them revise their manuscript.21
Such a use of AI could easily be generalised, and I would expect it to become a convention in the future given the strength of the technical skills AI brings to the assessment of scientific outputs.22
The breadth of AI models and their ability to relate any document to a large body of literature suggest to me that AI feedback could also be used to assess the contribution of a scientific paper. To make this judgement, one needs to assess its position in the existing literature, how it differs from it, what this difference adds in terms of insights, and how valuable these insights are. Such questions are where expert humans’ comparative advantage lies. That being said, this is also where subjective judgements might make humans most partial in their assessments. The relative imperviousness of AI models to these biases makes them an interesting complement to human reviews. In fact, human reviewers, knowing that their take will be complemented by an AI review, may end up submitting more disciplined and less partial reviews themselves.
Another reason to keep humans in the loop is the issue of accountability. Currently, the credit for accepting good papers and the blame for accepting bad ones fall on human editors. If reviews were fully automated, who would be responsible? AI companies? It is not impossible that some companies might be happy to take on such responsibility, but I suspect this is a major reason humans will retain final decision-making power in most scientific outlets.
AI pre-review
The next step is straightforward. If you know which model or tool will be used in the review process, submitters can use it themselves before submitting and get much of the feedback in advance. Refine already offers this directly to researchers. The AEA partnership currently uses the tool later in the editorial process and only for technical checking, but there is no technical reason why authors cannot perform the same type of check before submission.23
This starts to blur the distinction between writing and reviewing. Instead of waiting months for a referee to identify a problem, authors can expose every draft to something resembling a referee report while they are still writing it.
Equilibrium effects
With the advent of AI, it is hard to think that the old world of science will not be radically transformed. Even with only human writers, peer review already consumes a substantial amount of researchers’ time.24 With AI tools now dramatically accelerating the production of scientific papers, the pace at which papers can be produced is likely to become unworkable for the existing human peer-review system. AI review will come because it is better and because it is unavoidable.
As everybody changes their behaviour, one has to consider what new conventions will emerge in the game of scientific publishing, or, in other words, what the future equilibrium of that game will be. Overall, I am optimistic about the prospects offered by AI review.
Quality
First, I expect the end effect to be associated with an incredible increase in the quality of published research. AI is already reducing the cost of producing research. If researchers can write and analyse papers much faster, the number of submissions is likely to increase. But human attention remains limited, and therefore so does the number of papers that leading outlets can usefully publish. Competition for these scarce publication slots will become more intense, raising the quality threshold for publication. Instead of simply getting more papers like the ones we have now, we could end up primarily with much better papers.25
The gains in quality and rigour could be astonishing. For instance, it is hard to imagine some of the failures exposed during the replication crisis passing through peer review in the same way if every empirical paper were systematically checked for coding errors, underpowered designs, specification searching, weak identification and inconsistencies between its analysis and reported conclusions.
Speed
The turnaround could also become much faster. A remarkable amount of academic writing currently consists of anticipating reviewers’ criticisms: adding complementary analyses that a reviewer might request, expanding discussions to pre-empt possible objections. Part of the difficulty is the uncertainty. Since reviewers have varied perspectives, researchers have to try to figure out the distribution of possible objections their articles could face and address them in the manuscript.
AI pre-review has the benefit here of reducing this uncertainty by providing researchers upfront with the same type of feedback that reviewers and editors will see when they use the tool.
One fair concern here is a version of Goodhart’s law: when a measure of performance becomes a target, it eventually ceases to be a good measure because people learn to strategically optimise their score on that measure instead of focusing on good performance. In the case of AI reviews, people could learn to increase the score they receive from AI models by detecting some of the quirky factors the models pick up on when making a call. A recent study found, for instance, that strategically changing the presentation of a paper while keeping its scientific content unchanged could substantially increase AI review scores.26 Another more dramatic possibility is for people to try to cheat, for instance by inserting commands in their manuscript trying to trick the AI into giving a good review. A recent study found that this could indeed work.27
I feel, however, that such concerns are overblown. In a sense, you could say that Goodhart’s law already applies to scientific output, with researchers optimising for reviewers’ decisions perhaps more than for readers’ enlightenment.
Furthermore, I am sceptical that such questionable practices would become too frequent with AI reviews. The long record upon which scientists build their reputation means that they care a great deal about not staining this record. Many of the problems that emerged during the replication crisis and in cases of fraud arose in settings where the record could be obfuscated—for instance because data and code were not publicly available. A researcher who tried to buy their way to fame by hacking AI reviewers would have to gamble that whatever trick they found would not be identified and tracked down in past publications over the next 5 to 20 years. At the speed of technical progress, this would be a foolish gamble.
Content
A prominent role for AI in review would first reduce the need for what scientists call “marketing”. A lot of a paper’s writing is about selling it to the audience of potential reviewers. Authors carefully navigate the need to claim novelty without sounding too dismissive of past authors who are likely to review the paper. Part of it is also political: it involves citing, often positively, likely reviewers you do not want to offend, or people you think are amenable to your result and whom you would like to see selected. It can also involve not citing, or downplaying the citations of, people you do not want as reviewers.
Part of the writing is also about signalling your skills. I have described how economists, for instance, tend to signal with technical models while sociologists tend to signal with conceptual lingo. Signalling plays a role because of the incomplete information of reviewers. With a much greater ability to get to the bottom of the paper and much less propensity to be impressed by unnecessary technicality, the role of signalling might actually recede
Perhaps there will be a more radical change in the content of scientific papers. Papers are built the way they are because of the dynamics of persuasion between human brains. A paper makes a claim and supports it by describing the evidence for that claim, following clear rules so that reviewers or readers can understand and assess the hypotheses and the intermediate steps in the argument. Because scientists have an interest in overselling their results, a lot of academic writing is devoted to demonstrating as clearly as possible that the claim is not oversold and that possible concerns a reviewer might have are not vindicated.
An AI reviewer would not need all the same scaffolding to clarify the validity of the arguments made in a paper because it can much more quickly identify the exact nature of the claim, its degree and its domain of validity. It is therefore possible that papers will not require as many layers of defence. One of the most striking consequences is that AI may change, for instance, the way mathematical papers are written. AI might generate parts of papers in formal code largely unreadable by humans but directly checkable by formal verification systems, with AI handling much of the translation and checking. Relying on this verification, humans might now care less about the precise intermediary steps in the proof than about the results and their implications.
As a result, papers might become shorter and more direct, with less effort spent signalling sophistication or navigating disciplinary politics and more effort devoted to producing results that are clear, rigorous and checkable.
Criticisms of AI reviews are understandable given the limitations of earlier models, whose English writing ability was way ahead of their formal and technical abilities. Frontier models have changed this balance. Their technical abilities are now ahead of those of most experts. Scientists all over the world are experiencing, in an accelerated way, the intellectual demotion that elite chess players experienced from the late 1990s onward, when computers first defeated and then rapidly surpassed the world’s best human players. For experts who have dedicated their lives to mastering often extremely technical and arduous topics, this experience is humbling.
However, if we abstract from the feelings of scientists themselves, the strength of AI models should, in all likelihood, bring about an incredible scientific revolution in content and speed. Part of the revolution will be not just how scientific papers are produced but how they are vetted and selected. AI offers the prospect of dramatically improving a reviewing process which, in spite of its undeniable strengths, is riddled with biases and limitations.
Once high-quality assessment becomes cheap, immediate and reproducible, the distinction between writing a paper and having it reviewed becomes unclear. With AI tools, researchers will be able to continuously expose their work to the same kinds of scrutiny that previously arrived only months later from two or three randomly selected individuals.
That could eventually change not only peer review, but the speed, style and organisation of scientific research. Scientific articles could, for instance, focus on a summary of the results and implications for human readers, while technical proofs might be written in machine-verifiable code with AI handling much of the checking. I suspect that the current style of scientific contribution, based on often long-form papers, will likely be challenged and upstaged by new conventions better suited to AI’s ability to produce and check results.
References
Aczel, B., Szaszi, B. and Holcombe, A.O. (2021) ‘A billion-dollar donation: Estimating the cost of researchers’ time spent on peer review’, Research Integrity and Peer Review, 6, 14.
American Economic Association (2026) ‘AEA Journals Partner with Refine’, Annual Elections and Other Announcements, 14 August.
Azoulay, P., Fons-Rosen, C. and Graff Zivin, J.S. (2019) ‘Does science advance one funeral at a time?’, American Economic Review, 109(8), pp. 2889–2920.
Baldwin, M. (2017) ‘In referees we trust?’, Physics Today, 70(2), pp. 44–49.
Bornmann, L., Mutz, R. and Daniel, H.-D. (2010) ‘A reliability-generalization study of journal peer reviews: A multilevel meta-analysis of inter-rater reliability and its determinants’, PLOS ONE, 5(12), e14331.
Burnham, J.C. (1990) ‘The evolution of editorial peer review’, JAMA, 263(10), pp. 1323–1329.
Choi, B., Jun, T.J., Sung, J.W., Park, I.W., Lee, J.-M., Cho, S.I., Park, H.J., Lee, R.W. and Suh, J. (2026) ‘Invisible text injection and peer review by AI models’, JAMA Network Open, 9(1), e2552099.
Cornell, B. and Welch, I. (1996) ‘Culture, information, and screening discrimination’, Journal of Political Economy, 104(3), pp. 542–571.
Cortes, C. and Lawrence, N.D. (2021) ‘Inconsistency in conference peer review: Revisiting the 2014 NeurIPS experiment’, arXiv preprint arXiv:2109.09774.
Ellison, G. (2002) ‘The slowdown of the economics publishing process’, Journal of Political Economy, 110(5), pp. 947–993.
Erturk, M.S. and Durmaz, M. (2026) ‘Large language models as peer reviewers: Prompt sensitivity and model-dependent reproducibility’, Academic Radiology, 33(9), pp. 3642–3650.
Freeman, R.B., Xie, D., Zhang, H. and Zhou, H. (2024) ‘High and rising institutional concentration of award-winning economists’, paper presented at the NBER Science of Science Funding Initiative Conference, Cambridge, MA, 11 July.
Goodman, S.N., Berlin, J., Fletcher, S.W. and Fletcher, R.H. (1994) ‘Manuscript quality before and after peer review and editing at Annals of Internal Medicine’, Annals of Internal Medicine, 121(1), pp. 11–21.
Howell, A., Wang, J., Du, L., Melkers, J. and Shah, V. (2026) ‘Prestige over merit: An adapted audit of LLM bias in peer review’, arXiv preprint arXiv:2509.15122, revised 14 August 2026.
Huber, J., Inoua, S., Kerschbamer, R., König-Kersting, C., Palan, S. and Smith, V.L. (2022) ‘Nobel and novice: Author prominence affects peer review’, Proceedings of the National Academy of Sciences, 119(41), e2205779119.
Jefferson, T., Alderson, P., Wager, E. and Davidoff, F. (2002) ‘Effects of editorial peer review: A systematic review’, JAMA, 287(21), pp. 2784–2786.
Kennefick, D. (2005) ‘Einstein versus the Physical Review’, Physics Today, 58(9), pp. 43–48.
Kravitz, R.L., Franks, P., Feldman, M.D., Gerrity, M., Byrne, C. and Tierney, W.M. (2010) ‘Editorial peer reviewers’ recommendations at a general medical journal: Are they reliable and do editors care?’, PLOS ONE, 5(4), e10072.
Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., Chen, L., Ye, H., Liu, S., Huang, Z., McFarland, D.A. and Zou, J.Y. (2024) ‘Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI conference peer reviews’, in Proceedings of the 41st International Conference on Machine Learning, Proceedings of Machine Learning Research, 235, pp. 29575–29620.
Liu, Y., Sezgin, E. and Youngstrom, E.A. (2026) ‘Evaluating large language models for abstract evaluation tasks: An empirical study’, Frontiers in Research Metrics and Analytics, 11, 1807672.
Machado, D. (2026) ‘Generative AI bias against scientific novelty: A cautionary tale from a small-sample evaluation of research proposals’, Scientometrics.
Merton, R.K. (1968) ‘The Matthew effect in science’, Science, 159(3810), pp. 56–63.
Planck, M. (1949) Scientific Autobiography and Other Papers. Translated by F. Gaynor. New York: Philosophical Library.
Refine (2026) ‘Refine for Journals: State-of-the-art technical verification for peer review’.
Thakkar, N., Yuksekgonul, M., Silberg, J., Garg, A., Peng, N., Sha, F., Yu, R., Vondrick, C. and Zou, J. (2026) ‘A large-scale randomized study of large language model feedback in peer review’, Nature Machine Intelligence, 8, pp. 326–336.
Tomkins, A., Zhang, M. and Heavlin, W.D. (2017) ‘Reviewer bias in single- versus double-blind peer review’, Proceedings of the National Academy of Sciences, 114(48), pp. 12708–12713.
Uchida, H. (2025) ‘What do blind evaluations reveal? How discrimination shapes representation and quality’, working paper, revised 9 December 2025.
Yang, X., Sha, Z., Li, J., Yu, J., Sun, Y., Zhao, M., Fang, J., Guo, X., Wu, Y., Hu, X., Luo, Y., Liu, Q. and Wang, Z. (2026) ‘No hidden prompts needed! You can game AI peer review with presentation-only revisions’, arXiv preprint arXiv:2606.13044.
Cited in Kennefick (2005).
See Burnham (1990) and Baldwin (2017). Interestingly, the referee’s criticism was correct: Einstein later substantially revised the paper and reversed its central conclusion (Kennefick, 2005).
The evidence on how much peer review improves papers is thinner than one might expect. A classic before-and-after study of peer review and editing found a modest improvement in the reporting quality of accepted papers (Goodman et al., 1994). A systematic review concluded that evidence on the effects of editorial peer review remained limited (Jefferson et al., 2002).
My take might surprise some readers. It is based on my experience with frontier models, in particular ChatGPT 5.6 Sol. This model is significantly better than the previous 5.5, which was already very good. An interesting experiment by Paul Litvak found that 5.5 was able, on its own, to find 71 of 100 errors planted in psychology papers. I would expect this number to be even higher with 5.6. Unfortunately, Litvak’s benchmark did not feature a human benchmark against which to compare the performance of the AI models.
The slowdown of the publication process in economics, and the increasing role of extensive revisions, has long been documented (Ellison, 2002).
Planck (1949).
The prestige hierarchy is particularly concentrated in economics. Freeman et al. (2024) find that economics is the only one of 18 major academic fields with high and rising institutional concentration among award-winning researchers.
Merton (1968). The Matthew effect is named after the Parable of the Talents in the Gospel of Matthew in the New Testament, whose lesson is that “to the one who already has, more will be given.”
Huber et al. (2022).
See Uchida (2025). A similar field experiment at a major computer-science conference found that single-blind reviewers were more likely than double-blind reviewers to recommend papers from famous authors and top institutions (Tomkins, Zhang and Heavlin, 2017).
Howell et al. (2026).
Cornell and Welch (1996).
Kravitz et al. (2010).
In the study, there were more than 2 reviewers on average. But the coin-flip analogy still works. Around 72% of reviewers recommended a revision/acceptance. If reviewers’ judgments were entirely independent on a given paper (i.e. if they were like throwing coins with a 72% chance to land on accept/revise), the average proportion of disagreement in the paper sample would be 51% and the figure found in the article was 45.4%.
Cortes and Lawrence (2021).
A recent study by Erturk and Durmaz (2026) in radiology generated 720 reviews using eight different LLMs and found that agreement across models was somewhat higher than that typically observed among human reviewers. Inter-model agreement was κ = 0.25, compared with an average κ = 0.17 across human peer-review studies in Bornmann, Mutz and Daniel (2010), and κ = 0.08 among manuscripts evaluated by two reviewers in Kravitz et al. (2010). A separate study by Liu, Sezgin and Youngstrom (2026) found moderate agreement between LLMs and human reviewers on more objective aspects of papers, such as clarity, objectives and results, while agreement was much weaker on more qualitative judgements such as impact, engagement and applicability.
For the statistician reader, I tend to see AI models trained on large datasets as similar to non-parametric estimators, whose predictions in sparsely represented regions are typically biased and pulled towards more common patterns in the data. Similarly, an AI model faced with a radically new idea might struggle to share the author’s enthusiasm and instead reproduce the critical takes that people not as specialised in the topic as the author would have about it on average.
Machado (2026).
Liang et al. (2024).
Thakkar et al. (2026).
American Economic Association (2026).
One concern that is legitimately raised about the use of AI by reviewers is confidentiality: manuscripts that are uploaded onto language models can be used for training (if the LLM settings do not exclude it), and the intellectual content of the paper could potentially be taken and spread to others before the publication of the material and its proper attribution to the researchers who produced it. This is not an argument against AI per se, but against sharing confidential intellectual content with AI models without appropriate safeguards. Secure AI review that excludes the use of material for training purposes could be organised by scientific journals and conferences. Some academic publishers already allow limited use of AI tools by reviewers where confidentiality is protected.
Refine (2026).
Aczel, Szaszi and Holcombe (2021) estimate that active reviewers completed about 4.7 reviews a year, at roughly six hours per review, or three to four working days per reviewer each year. This work is, however, concentrated among a minority of reviewers.
Joshua Gans made the same point in a post earlier this year.
Yang et al. (2026).
Choi et al. (2026).














