We ran the academic-peer-review skill blinded on the original submitted version of an open-review article, then scored it against the published referee reports and the revisions the authors actually made — and we publish what it missed, not just what it caught.
The skill was executed blinded on the original submitted version (v1) of the article — the agent was barred from fetching the article's pages, its reviews, or anything about it online. Separately, the published referee reports and the authors' v1→v2 revision notes were harvested. A comparison pass coded every human and skill comment with the skill's own 20-class taxonomy and matched them conservatively (doubtful matches resolved against the skill). An independent adversarial audit then re-verified the highest-impact claimed matches against the underlying texts; its corrections are disclosed per article below.
The retained PeerJ article was outside the skill's n=60 calibration corpus and was published in March 2026. Publication timing alone does not establish absence of model contamination. PeerJ's review history is published under CC BY and quoted with attribution. The withdrawn F1000Research arm is excluded from the retained results.
Runtime: Claude (claude-fable-5) running the skill as an agent. A note on reading the numbers: published referee reports differ enormously in granularity — one PeerJ reviewer alone contributed a long list of bench-protocol questions — so “comments caught” is not comparable across articles. The measures that matter most are whether the skill found the defects that decided the outcome, whether its verdict matched, and what it uniquely added.
Optimizing in vitro osteoclastogenesis: bone marrow-derived macrophages differentiation and cell density as critical determinants — PeerJ (20995, submitted 2025-09-23, published 2026-03-25)
Adversarial audit: 2 of the 6 highest-impact claimed matches downgraded to same-area-only (the skill flagged reporting completeness where the human flagged test validity, and critiqued the opposite direction of one claim), and one novel finding partially overlaps a human comment. The three headline novels survived the audit.
On the match count: 6 of 49 at row level (the run's own summary claimed 12). The comparison matrix marks 6 rows found by both — 3 exact, 3 partial — and the adversarial audit downgraded 2 of those to same-area-only, leaving 4 substantive matches.
Review history (4 referees + editor decision) published at peerj.com/articles/20995/reviews under CC BY.
A second arm (F1000Research 15-579, clinical) was published here until 2026-09-06 and has been withdrawn. The v1 PDF supplied to the skill was the journal's own article PDF, which appends the published referee reports and author responses on pages 11–19 — the very reports the run was scored against. The skill's own run notes disclosed this; this page did not. It will be re-run on the manuscript body alone and republished with that provenance stated.
The skill found the outcome-deciding defect, matched the editor's verdict, demanded 5 of the 9 review-driven revisions, and surfaced 3 audited findings no human reviewer raised. Its consistent weakness is breadth on hands-on bench detail — reagent-level protocol questions and figure-formatting conventions that an experienced wet-lab referee accumulates by habit. It is a strong reviewing partner and screener; it is not a replacement for a domain expert's eye, and the skill itself says so.