Abstract
Comprehensive basketball video understanding requires resolving not only what event occurs, but also who is responsible and when the key evidence appears. However, existing methods typically treat spatial perception and semantic recognition as isolated tasks, failing to ground events to individual players or pinpoint their temporal boundaries within complex collective dynamics. To bridge this gap, we introduce BasketEvent, a player-centric basketball event understanding dataset curated from real NBA broadcasts. In BasketEvent, event labels are grounded to the responsible players, and a manually annotated subset of 998 samples with precise event intervals is provided to evaluate temporal evidence localization. Based on this data, we propose PlayNet, a player-centric reasoning framework that maps basketball videos to player-level event predictions with temporal evidence. Concretely, PlayNet tracks key entities, associates player identities, and reasons about events by modeling player-player, player-ball, and global court interactions, while aggregating sparse temporal evidence via gated pooling. Extensive experiments demonstrate that PlayNet significantly outperforms representative video-level and crop-based baselines, proving the superiority of player-centric modeling for fine-grained sports video understanding. Our data, code, and models will be made publicly available.
Dataset
BasketEvent is a player-centric basketball event understanding dataset curated from real NBA broadcasts. Starting from official NBA play-by-play records, we collect event-centered video clips and normalize structured metadata into 10 representative event categories: Missed Shot, Made Shot, Free Throw, Foul, Turnover, Jump Ball, Rebound, Steal, Block, and Assist. Each event is grounded to the responsible player, enabling models to reason about not only what happens, but also who performs the event.
The dataset contains approximately 35K videos, 90.6 hours of broadcast footage, and 51K player-level samples. To evaluate generalization, BasketEvent is split at the game level, with 189 games for training and 37 unseen games for testing. We further provide a manually annotated temporal evaluation subset of 1,000 player-level samples, where each sample includes precise start and end timestamps for the target event, supporting evaluation of when the key evidence appears.
Method
We propose PlayNet, a player-centric framework that predicts who-what-when event tuples from basketball videos. PlayNet grounds players and the ball, associates player identities, and extracts trajectory-guided features for each visual entity.
The model reasons over global-player, player-player, and player-ball interactions, then uses gated temporal pooling to focus on decisive moments and localize when the event evidence occurs.
Main Results
Player-level Event Recognition
PlayNet consistently achieves the best performance across all recognition metrics. Compared with the strongest interaction-aware baseline, GroupFormer, PlayNet improves the macro F1-score from 0.555 to 0.669 and Rec@1 from 0.572 to 0.700. It also reaches 0.888 Rec@3 and 0.953 Rec@5, demonstrating the effectiveness of jointly modeling global court context, player-player interactions, and player-ball dynamics.
Temporal Evidence Grounding
On the manually annotated temporal subset, PlayNet achieves end-to-end Hit@0.1/0.3/0.5 scores of 0.448/0.284/0.138, clearly outperforming the training-free selectors and zero-shot VLMs. When evaluated only on correctly recognized events, it obtains TP-only Hit@0.1/0.3/0.5 scores of 0.608/0.385/0.187. These results show that the learned clip-level gates capture meaningful coarse temporal evidence without explicit temporal-boundary supervision, while precise localization remains challenging.
Ablation Study
Component Ablation
Removing key components consistently hurts recognition performance. Compared with the simple ROI+Gate baseline, full PlayNet improves F1-score from 0.498 to 0.669, showing the importance of global context, player-player interaction, and ball dynamics.
Sampling Ablation
We study the effect of frame rate and clip number. The setting FPS=4 and M=12 gives the best overall recognition result, balancing motion detail and temporal coverage.
Qualitative Examples
These examples show PlayNet predictions on BasketEvent videos. For each case, the model predicts who performs the event, what event occurs, and when the supporting evidence appears.