No single human-level verdict: LLM reviewers overlap with humans on broad feedback but catch under half of real flaws and rarely reject
Asked (summary)
What empirical studies have compared LLM-generated scientific peer review with human peer review, and what did each study find about agreement, error detection, recall, score bias, overconfidence, and limitations?
This report covers 27 empirical studies published June 2023 to September 2026 — a curated subset of a 136-row evidence corpus drawn from 41 independent source hosts — spanning direct human–LLM review comparisons and controlled planted-error benchmarks. A blank cell means the study did not measure or report that dimension, not that no effect exists. Several results are arXiv preprints; publication status is marked per study.
Click any cell for the study's full reported result on that dimension. Sorted by year.
quantified or detailed findingunfavourable to LLM (bias, misses, hallucination)comparable to or above human baseline— not measured / not reported
Chronology
Studies by year; colour marks direct human–LLM comparisons vs adjacent controlled benchmarks.
What the matrix cannot say alone
Agreement is task-dependent, not a verdict: GPT-4's comment overlap with an individual human reviewer was 30.85% on 3,096 Nature-family papers and 39.23% on 1,709 ICLR papers — comparable to human–human overlap of roughly 28–35% — which signals complementarity, not equivalence. ar5iv.labs.arxiv.org
Controlled flaw detection stays under half: early GPT-4 found 7 of 13 inserted errors, but PaperAudit-Bench's best error coverage was 51.4% (GPT-5, best macro-F1 ≈41%) and no model on FLAWS exceeded 50% even with 10 candidate answers (best 39.1%). arxiv.org
Scores rarely track human judgment: PaperAudit's Spearman correlation with human scores was 0.142; in a 300-manuscript ophthalmology comparison humans rejected 73.33% while ChatGPT-4o and Gemini 1.5 each rejected only 2% — acceptance bias, not agreement. ovid.com
LLM reviewers can favour LLM work and hallucinate confidently: LLM-REVal's simulated reviewers scored LLM-authored papers 6.2142 vs 5.9371 for human papers (acceptance 78% vs 49%) despite r=0.5046 validation correlation with human scores, and PaperAudit reported 46.8% AI-extra major points beyond human reviews. arxiv.org
Practical synthesis for editors and researchers
Use LLM review as triage and augmentation with expert verification, never autonomous replacement: recall rises for salient, commonly raised issues, but misses of real flaws remain substantial in every controlled benchmark.
Do not read fluent, specific prose as calibrated correctness: unsupported extra critiques, false alarms, and hallucinated content are documented across ReviewerGPT, PaperAudit, MARG and the multimodal audit.
Human review is itself an incomplete gold standard: an LLM-only concern may be a valid novel point or a hallucination, and only expert adjudication separates them; recall against human reviews undercounts novel critiques while overlap metrics reward imitation.
Expect sensitivity to prompts, model versions, paper length, author identity cues, sycophancy and prompt injection; most samples are computer-science-heavy and contamination-prone, so results transfer poorly across venues and fields.
All 27 studies
Study
Year
Type
Status
Design / sample
Source
Method: 27 primary empirical studies selected editorially from a 136-row assembled evidence corpus (41 source hosts; no single URL backs more than half the rows), current through 23 September 2026. Matrix cells quote each study's own reported result; blanks mean not measured or not reported. Dimensions: agreement/overlap with humans, objective flaw detection, recall of human concerns, score or accept/reject bias, overconfidence/false positives, and stated limitations. Roughly 100 adjacent rows — secondary pages, system-building papers without direct human comparison, and duplicates — were cut for space. Numbers are as published; preprints are unrefereed.