No single human-level verdict: LLM reviewers overlap with humans on broad feedback but catch under half of real flaws and rarely reject

Asked (summary)

What empirical studies have compared LLM-generated scientific peer review with human peer review, and what did each study find about agreement, error detection, recall, score bias, overconfidence, and limitations?

This report covers 27 empirical studies published June 2023 to September 2026 — a curated subset of a 136-row evidence corpus drawn from 41 independent source hosts — spanning direct human–LLM review comparisons and controlled planted-error benchmarks. A blank cell means the study did not measure or report that dimension, not that no effect exists. Several results are arXiv preprints; publication status is marked per study.

Evidence matrix: 27 studies × 6 finding dimensions

Click any cell for the study's full reported result on that dimension. Sorted by year.

quantified or detailed finding unfavourable to LLM (bias, misses, hallucination) comparable to or above human baseline — not measured / not reported

Chronology

Studies by year; colour marks direct human–LLM comparisons vs adjacent controlled benchmarks.

What the matrix cannot say alone

Agreement is task-dependent, not a verdict: GPT-4's comment overlap with an individual human reviewer was 30.85% on 3,096 Nature-family papers and 39.23% on 1,709 ICLR papers — comparable to human–human overlap of roughly 28–35% — which signals complementarity, not equivalence. ar5iv.labs.arxiv.org
Controlled flaw detection stays under half: early GPT-4 found 7 of 13 inserted errors, but PaperAudit-Bench's best error coverage was 51.4% (GPT-5, best macro-F1 ≈41%) and no model on FLAWS exceeded 50% even with 10 candidate answers (best 39.1%). arxiv.org
Scores rarely track human judgment: PaperAudit's Spearman correlation with human scores was 0.142; in a 300-manuscript ophthalmology comparison humans rejected 73.33% while ChatGPT-4o and Gemini 1.5 each rejected only 2% — acceptance bias, not agreement. ovid.com
LLM reviewers can favour LLM work and hallucinate confidently: LLM-REVal's simulated reviewers scored LLM-authored papers 6.2142 vs 5.9371 for human papers (acceptance 78% vs 49%) despite r=0.5046 validation correlation with human scores, and PaperAudit reported 46.8% AI-extra major points beyond human reviews. arxiv.org

Practical synthesis for editors and researchers

All 27 studies

StudyYearTypeStatusDesign / sampleSource

Method: 27 primary empirical studies selected editorially from a 136-row assembled evidence corpus (41 source hosts; no single URL backs more than half the rows), current through 23 September 2026. Matrix cells quote each study's own reported result; blanks mean not measured or not reported. Dimensions: agreement/overlap with humans, objective flaw detection, recall of human concerns, score or accept/reject bias, overconfidence/false positives, and stated limitations. Roughly 100 adjacent rows — secondary pages, system-building papers without direct human comparison, and duplicates — were cut for space. Numbers are as published; preprints are unrefereed.

This report was generated automatically by Keenable SELECT at a user's request, from publicly available web sources linked herein. Keenable does not review, verify, or endorse its contents and makes no representation as to accuracy, completeness, or timeliness; AI-based extraction may contain errors. Nothing in this report is investment, legal, financial, or other professional advice. All trademarks and referenced content remain the property of their respective owners; no affiliation or endorsement is implied. To report an error, rights concern, or request removal: legal@keenable.ai.

Keenable SELECTAsk your own question
Made with Keenable SELECT