AI peer review goes pro
A what we're reading spotlight
Last week Refine announced partnerships with the AEA and the Econometric Society, both of which will now use its AI-assisted technical review in their publication process. The current implementation is narrow, only a “pre-publication audit of technical execution and deposition” to raise issues that would rise to the level of a correction if caught post-publication. At this stage, it’s not being used as an input for deciding what gets published (and Refine’s “overall feedback” is not shown to editors). Zarek Brot, who was in the AEA’s pilot, has a useful thread on the author experience.
Predictably, there have been a lot of takes.
On the pro side, the arguments I find most compelling are: (a) Peer review has been buckling under its own weight for years, with editors even at top journals struggling to find reviewers. That problem gets much worse if AI drives a substantial increase in submissions, as it seems to be doing. Something’s got to give, and AI is likely to be part of the answer, so I would much rather it be via well-designed and publicly-validated tools than reviewers pasting manuscripts into a chat (note: we covered Refine’s own evaluation a few weeks back, and Paul Litvak has a new AI review eval that I will discuss more in a future post).
And (b) There are several subtasks that human reviewers have never been especially good at — like catching math or software errors, checking consistency between tables and text, and verifying that citations accurately represent cited work — but that we’ve fully relied on them for nonetheless. Offloading those subtasks could let reviewers concentrate on the tacit and taste-based components of review, like assessing the novelty of results, their impact on the field, how they fit (or don’t) with existing literature, etc. Peter Hull, Paul Novosad, and Kevin Bryan each proposed related versions of what the future of combined AI and human review could look like.
On the con side: (a) At equilibrium, there will need to be ways to avoid creating capture for for-profit companies – i.e. if journals pay subscription fees for AI review on one end, that heavily incentivizes authors planning to submit to those journals to pay for those services up front as well. That is less of an issue for this pilot, where Refine’s use does not affect publication decisions and is being offered to journals for free, but if it ever shifts upstream then this concern would need to be resolved.
And (b) There is also a possibility that offloading technical checks could make human reviewers less careful/thorough, which could reduce overall review quality on net. This is something Refine and the journals could formally study, and I hope they do.
Overall, props to Ben Golub for his work on Refine, and to the AEA and the Econometric Society – two of the publishers who have also been on the front lines of replication packages and data checks – for being forward-looking and experimental about research evaluation.


