AI and academic publication
Three new working papers make predictions about how AI will affect the price of submissions, the review system, and the size of research teams...
Over the last two months, I have accepted almost every invitation to review academic papers. On almost every one, I recommended rejection. Nearly all had been written with a large language model, which is fine by me, since I, too, use Claude Code or Codex for every paper I write. The trouble was that they were terrible in the same way, with an abstract opening on a question, an identical section order, irrelevant and old literature, and so much unnecessary mathematics (even in econ history).
In fact, I was not capable of following the math in most of the papers, so my own model would check the algebra. In almost all cases the result reported in the text was the opposite of the result the model had derived. Few of the authors had taken the trouble to actually check – themselves or with a different model – whether the math actually works out. The worst was a submission to a well-known journal with a section headed ‘placeholder results’, explaining that the next step was to download the data and run the analysis. The three figures were blank. The authors had not even taken the trouble to check their own paper for completeness. (The editor should have caught it, sure, but editors are drowning too, so I have sympathy for that oversight.)
Which is why the news of 6 August was good news. Refine – an AI reviewing tool built by the economists Ben Golub and Yann Calvó López, and one I use far too heavily – announced partnerships with the American Economic Association and the Econometric Society. Both now run its technical verification before publication, while mistakes are still cheap to correct. Earlier this year, John Cochrane predicted it:
Most referee reports do not identify the major point of the paper, and do not assess if the paper backs up that point. They do not notice glaring gaps of logic, basic theorems violated, econometrics advice 101 ignored. Editors are lucky if one out of three reports is vaguely useful. Clearly, this task is going to be radically impacted by AI. If I were an editor, I’d feed every paper to refine on receipt. I will surely get refine’s opinion before any referee report I write in the future.
But the reaction on X was less settled. Joshua Gans thought it belonged at submission (as Cochrane had proposed), where it can still change the paper; others asked whether we want a private toll on submission at all. (Zachary Horvitz found that renaming a file to ‘paper_final_draft_ready_for_review.pdf’ raises its score, which tells you roughly where we are.) And behind it sits a more radical (hopeful) claim, made, for example, by Adam Mastroianni: the journals are finished, and publishers storing PDFs at a forty per cent margin is unconscionable.
The pathology is real but the prognosis is wrong. Journals are about to become more important, not less, and having spent two months writing three working papers on the question, let me say why.
Since ChatGPT, new economics papers on arXiv have almost doubled, while NBER working papers, where posting is gated by affiliation, have risen about fifteen per cent. But there is no more expert attention this year than last, and there will be no more next year. Economics rations that attention by making us wait, which averages more than two years to acceptance and buys the journal no extra capacity. A fee would ration it at entry instead, and could pay the referee for the reading. So how many journals actually charge? That is what the first paper, ‘The Price of Submission’, set out to measure. In July I collected the fee and referee-payment policy of the 500 highest-ranked economics journals. (Claude agents did the collecting for a few dollars, a small demonstration of the thing the paper is about.)
One journal in five charges a fee, rising from 2 per cent of the lowest-ranked fifty journals to 68 per cent of the top fifty, at a median of $143. Now ask a different question: what price would actually clear the queue? The paper will not be pinned to a dollar – it answers in shares of the prize a placement is worth – so the figures here are my own back-of-the-envelope: about $412 at the top of the ranking today, and perhaps $1,555 in five years if capability keeps climbing, with the referee wage that balances the books rising from $206 to $799 a report. Those are mine, not the paper’s. They measure pressure on the queue rather than anyone’s price list, and the gap between the $143 charged and the $412 I get is the pressure. What I would watch is the line above which journals charge authors and pay referees, and below which they sell open-access publication instead. On the fees we can actually compare with 2019, most have gone up. And since a $1,555 fee reads differently from Stellenbosch than from Stanford, the eighty-eight journals already waiving fees for poorer countries matter.
If a machine can do the checking, what is a journal for? That is the second paper, ‘The Value of Peer Review and the Reward to Reputation’. A journal reads carefully what it can and deals with the overflow in some other way. In the past, that meant rejecting papers unread; now it can mean letting a machine read them. Both routes produce the same public stamp, but here the intuition becomes less obvious. Every paper the journal declines to certify joins the pool of uncertified papers against which that stamp is valued. So when a journal rejects good papers unread, it improves the quality of the uncertified pool and thereby weakens the value of its own certificate. A cleaner certificate can therefore be worth less than a dirtier one. At the paper’s illustrative parameters, the two rules cross exactly once, when about two submissions in five are good.
The result that matters for Refine is sharper still. Error rates on a codifiable screen end up reflecting our own behaviour. Once everyone has access to the same tool, authors will run it before submitting and fix whatever it flags. That is exactly the point the economist Luis Garicano made when the partnerships were announced: papers will already have passed through the AI referee, so it will have little new to find. At a selective journal, where few submissions clear the bar, this kind of repair roughly halves the value of a machine-screened stamp. Rejection unread, however, survives at any repair cost, and so does human judgement. Expert judgement remains the only test an author cannot rehearse in advance. The same is true of a grant panel or a hiring committee.
The third paper, ‘AI and the Research Team’, asks what happens to us. Think of a laboratory as one leader’s judgement combined with execution that can be scaled up. Better AI then pulls in two directions. It makes execution cheaper, so the laboratory grows, but it also automates codifiable tasks, so each unit of execution requires fewer people. Put those effects together and team size can turn at most once. In some fields it rises and then falls; in others it simply keeps rising. What determines the difference is how much of the work remains irreducibly human. Where none does, the team eventually shrinks to one person. Perhaps that is the least intuitive result in the three papers: the fields where machines help most may end up with the fewest scholars in the room.
Which fields? I scored 3,600 abstracts across 30 fields for how much physical execution they require, fixing the scores before looking at any outcomes. Seven of the thirty fall below the threshold. Mapping the METR capability trend onto research tasks then gives a date for the turn: 2027 is the last year in which research teams grow in algebra and number theory and in theoretical computer science, or 2028 if you allow for the slow adoption of new tools. Economics clears the threshold by just 0.001, which is close enough to a coin landing on its edge that we should be cautious. None of this is confirmed yet. In a panel of 2,279 principal investigators running to the end of 2025, the estimated effect is 0.19 with a standard error of 0.26. That is roughly what we should expect from data that end before the predicted transition begins.
The same divide runs through all three papers. AI has made producing research very cheap. It has not made checking research cheap, or judging it, or deciding which question is worth asking. Mastroianni is right that publishers extract too much for too little, but I think he draws the wrong conclusion. The scarce input in science is increasingly what a good journal can still provide: someone who has read the paper carefully and is willing to say what they make of it. There will be fewer such people, they will be paid, and their judgement will become more valuable.
Which brings me back to the paper with the blank figures. Refine.ink would have caught the problem in seconds, and the paper should have been stopped long before it reached me. But someone still had to decide that the paper should not exist. No machine made that judgement. I did.







