AI-Written Research Proposals Rated Similarly to Human Ones, But AI Reviewers Show Systematic Pro-AI Bias

Large language models can produce scientific project plans comparable in quality to those written by human researchers, but AI-based evaluators show a systematic preference for AI-generated proposals, according to a new study published on arXiv.

The study, led by Jia Liu of the Kavli Institute for the Physics and Mathematics of the Universe, compared one-page research proposals written by humans and by three LLMs (ChatGPT 4o, Claude Sonnet 4, and DeepSeek V3) across eight expert-conceived projects in physics, astrophysics, and cosmology.

Eight domain experts each outlined a research project covering topics including active galactic nuclei evolution, the galaxy-dark matter halo connection, weak lensing intrinsic alignments, gravitational-wave black hole binaries, and tests of general relativity with pulsar timing arrays. Human project planners, graduate students or postdocs expert in each topic, then wrote one proposal each without AI assistance. Separately, undergraduate students outside the specific fields used a fixed prompt template to generate three AI proposals per project, producing a total corpus of 32 proposals: eight human-written and 24 AI-generated.

Four senior researchers served as blind human reviewers, evaluating proposals on clarity, appropriateness of methods, resource planning, and feasibility. Two frontier LLMs, Claude Opus 4.8 and ChatGPT Pro 5.5, performed the same evaluation.

We believe news should be guided by evidence, not sensationalism. Your support helps make that possible.

Support independent reporting

Human reviewers rated human- and AI-written proposals similarly overall, with a mean score of approximately 3.5 out of 5 for both groups. However, AI reviewers scored AI-generated proposals roughly one point higher than human-written ones, revealing a systematic pro-AI bias. This bias was concentrated in projects where the human-written plans were weaker; the correlation between human plan quality and AI reviewer bias was approximately -0.95.

When asked to identify whether each proposal was human- or AI-written, human reviewers achieved 72 percent accuracy for human-written proposals and 79 percent for AI-written ones, with individual accuracy ranging from 59 to 88 percent. Human reviewers identified AI-written proposals by characteristics such as template-like five-step methodology sections, absence of citations, round-number timelines, and overuse of machine-learning terminology. Human-written proposals were distinguished by specific author-year citations, idiosyncratic domain jargon, first-person phrasing, and focused scope.

Both AI reviewers correctly classified all 32 proposals with 100 percent accuracy, applying the same cues consistently and without error, citing features such as uniform five-phase templates, polished impersonal prose, generic survey choices, exhaustive tool lists, and boilerplate risk mitigation.

The paper notes that the results suggest caution when deploying LLMs widely in proposal preparation and evaluation, as AI evaluators may introduce systematic bias in favor of AI-generated content.

Images: (none; research paper)

Sources: arXiv:2607.25881 (Liu et al., Jul 28, 2026)

Scroll to Top