
The comforting story about artificial intelligence in science holds that automation will remove the drudgery and researchers will use the reclaimed time to think more deeply and finish their work more completely. A new modeling study argues this story gets the economics exactly backwards. By making scientists faster, large language models raise the price of a researcher’s time, and time that is more expensive gets spent on the next project, not on polishing the current one. The authors, Eamon Duede, Kevin Gross, Molly Crockett, and Carl Bergstrom, predicted bluntly in their abstract that researchers will do more, less well, rather than the same amount, better.
The paper, posted to arXiv on 19 July 2026 and not yet peer reviewed, is deliberately generous to the technology. The model assumes LLMs work exactly as their most enthusiastic proponents imagine: cheap, fast, accurate, and free of errors. Even under that assumption, the authors show, the mere existence of the tools reshapes how scientists allocate effort, because it changes the frictions and incentives that steer research. The finding is structural, not a complaint about hallucination or deskilling, and the authors are explicit that the AI tools are not the root problem. LLMs hold up a mirror to an incentive system that already rewards quantity over quality.
The mathematical machinery comes from behavioral ecology. The model borrows optimal foraging theory, developed by Eric Charnov in 1976 to describe how an animal decides when to leave a patch of food, and applies it to a researcher deciding when to leave a project. A scientist spends a discovery phase of fixed length learning a project’s value: hypotheses, pilot experiments, preliminary data. Once the value is known, the researcher either abandons the project and starts a new discovery, or invests in development, which has two parts. The first is required work, the irreducible cost of turning results into a publishable paper: figures, typesetting, proofreading, submission. The second is discretionary development: follow-up experiments, sensitivity analyses, deeper statistics, polishing prose. The model’s key quantity is the long-run rate of return on time, which sets the opportunity cost of every hour a researcher spends.
From this framework the paper derives three predictions, and the first two are uncomfortable. When LLMs shorten the discovery phase, as they do in technical fields where they help identify promising problems quickly, researchers become more selective, publishing a smaller fraction of projects, but even the projects they do publish get developed less thoroughly, because the opportunity cost of time has risen. When LLMs instead reduce the required work of producing a manuscript, which is the case in fieldwork-based disciplines where writing and analysis dominate, the threshold for publishing falls, so more projects get published, each developed less thoroughly. Only the third case is unambiguously good: when LLMs accelerate discretionary development itself, making follow-up experiments and deeper analysis cheaper, projects get developed more thoroughly. Which outcome dominates in a given field depends on which phase of research the LLM speeds up.
The model’s most pointed result is a diagnosis of the time-savings fallacy. The hope that automating mundane tasks will buy scientists time to think deeply fails, the authors write, because LLMs make researcher time more valuable, not less. Once a paper is nearly publishable, the marginal hour spent refining it is an hour not spent discovering the next project, and when everything moves faster, that hour’s opportunity cost grows. The result is an incentive to move on. The authors note that they have treated scientific institutions as static, which is reasonable on short timescales but guarantees a period of mismatch as practices change faster than policies. Journal submissions are already rising sharply in fields where LLMs cut the time to produce a manuscript, straining peer review, and the incentive structure of publish or perish predates the tools by decades.
The paper’s most striking detail is its own AI use statement. The authors report that the text was written without AI assistance beyond built-in spell checking, but that they used ChatGPT 5.4, ChatGPT 5.5, and Gemini 3.1 to suggest proof strategies, Claude Fable 5 to check proofs for consistency, and ChatGPT 5.5 to refine figure fonts and LaTeX formatting. A paper arguing that LLMs will push scientists toward haste was itself produced with LLM assistance at exactly the stages the model describes: the discretionary work of proof-checking and formatting, accelerated and partially delegated. It is a small, honest illustration of the phenomenon rather than a critique of it.
The caveats are real. This is a preprint, and the model simplifies heavily, treating time as the only resource and holding incentives fixed. Real LLMs make errors, and their errors may impose costs the model ignores. But the value of the exercise is that it isolates the structural effect: even a perfect tool, in a system that rewards output, shifts effort away from depth. The practical conclusion for science policy is that fixing the outcome means changing the incentives, because the technology is not the binding constraint. The authors argue that the hope that time savings will translate into deeper thinking is tempered by their results. The mirror the LLMs hold up is showing the system its own reflection.
References
Eamon Duede, Kevin Gross, M.J. Crockett, and Carl T. Bergstrom, The unintended consequences of large language models as a labor-augmenting technology in science. arXiv:2607.17397 (2026). DOI: 10.48550/arXiv.2607.17397.
Kaia Glickman, Scientists using LLMs will ‘do more, less well’, modelling study predicts. Nature News, 31 July 2026. DOI: 10.1038/d41586-026-02397-5.
E.L. Charnov, Optimal foraging, the marginal value theorem. Theoretical Population Biology 9(2), 129-136 (1976).

