Learning Compact Spatio-Temporal Representations for Edge Video Keyframe Selection Using Masked Graph Autoencoders
DOI:
https://doi.org/10.64972/dea.2023.v2i3.3704d:40-51Keywords:
Masked Graph Autoencoder, Few-Shot Fusion, Edge Video Analytics, Keyframe Selection, Spatio-Temporal GraphAbstract
Selection of keyframes for edge video analytics needs to retain rare but operationally significant events within strict computation and bandwidth limitations. This paper proposes a Masked Graph Autoencoder for edge video keyframe selection with few-shot fusion. Video clips are converted into sparse spatio-temporal graphs where nodes contain frame-level appearance, motion and object-interaction descriptors, and adaptive edges encode temporal continuity and semantic neighborhood relations. Add a masked reconstruction objective and prototype-guided few-shot fusion to have the selector learn from a small number of labelled samples without overfitting to the dominant background frames. Experiments in an edge-oriented benchmark for traffic, campus, retail and indoor monitoring scenarios show that the proposed model has an F1 score of 88.7%, reduces redundant selected frames by 31.6%, and lowers the average per-clip decision latency to 18.4ms on an embedded GPU. Lightweight transformers, temporal clustering and graph autoencoder baselines have been used for comparison; under 5-shot supervision and at 20% packet loss, the method improved mean average precision by 5.9-11.8 percentage points and remained stable. Based on the above results, masked graph reconstruction and few-shot fusion provide a feasible path for high-fidelity, compact and deployable keyframe selection in distributed edge video systems.
Downloads
Published
How to Cite
Issue
Section
License
Copyright (c) 2023 Klementa Hronová, Jaroslav Skála, Branko Malina

This work is licensed under a Creative Commons Attribution-NonCommercial-ShareAlike 4.0 International License.