Action progress prediction, the task of estimating how far an action has advanced toward its completion, is a fundamental yet underexplored problem in video understanding. Accurate progress estimation has broad implications for augmented reality assistance, procedural monitoring, collaborative robotics, and error detection in complex sequential tasks. Building upon the ProgressNet architecture, this work extends prior approaches along multiple directions. First, a phase-based progress annotation strategy is introduced for the Mobile ALOHA dataset, enabling the generation of semantically meaningful supervision signals. Second, a comprehensive evaluation is conducted across four heterogeneous benchmarks, Cholec80, Mobile ALOHA, MECCANO, and EGTEA Gaze+, spanning robotic manipulation, surgical procedures, and unconstrained first person activities. We systematically investigate the impact of data augmentation strategies, feature extractor choices, and alternative loss functions, including a Boundary Observant loss and a novel regularization term designed to enforce temporal consistency. Finally, we assess model robustness through controlled temporal distortions demonstrating that the learned representations rely predominantly on visual content rather than implicit. Overall, this work highlights the importance of structured temporal representations and tailored training objectives for accurate action progress estimation, and provides insights into design ing models that generalize across diverse domains and temporal conditions.
Action progress prediction, the task of estimating how far an action has advanced toward its completion, is a fundamental yet underexplored problem in video understanding. Accurate progress estimation has broad implications for augmented reality assistance, procedural monitoring, collaborative robotics, and error detection in complex sequential tasks. Building upon the ProgressNet architecture, this work extends prior approaches along multiple directions. First, a phase-based progress annotation strategy is introduced for the Mobile ALOHA dataset, enabling the generation of semantically meaningful supervision signals. Second, a comprehensive evaluation is conducted across four heterogeneous benchmarks, Cholec80, Mobile ALOHA, MECCANO, and EGTEA Gaze+, spanning robotic manipulation, surgical procedures, and unconstrained first person activities. We systematically investigate the impact of data augmentation strategies, feature extractor choices, and alternative loss functions, including a Boundary Observant loss and a novel regularization term designed to enforce temporal consistency. Finally, we assess model robustness through controlled temporal distortions demonstrating that the learned representations rely predominantly on visual content rather than implicit. Overall, this work highlights the importance of structured temporal representations and tailored training objectives for accurate action progress estimation, and provides insights into design ing models that generalize across diverse domains and temporal conditions.
Action progress prediction in videos
MATTESCO, ANTONIO
2025/2026
Abstract
Action progress prediction, the task of estimating how far an action has advanced toward its completion, is a fundamental yet underexplored problem in video understanding. Accurate progress estimation has broad implications for augmented reality assistance, procedural monitoring, collaborative robotics, and error detection in complex sequential tasks. Building upon the ProgressNet architecture, this work extends prior approaches along multiple directions. First, a phase-based progress annotation strategy is introduced for the Mobile ALOHA dataset, enabling the generation of semantically meaningful supervision signals. Second, a comprehensive evaluation is conducted across four heterogeneous benchmarks, Cholec80, Mobile ALOHA, MECCANO, and EGTEA Gaze+, spanning robotic manipulation, surgical procedures, and unconstrained first person activities. We systematically investigate the impact of data augmentation strategies, feature extractor choices, and alternative loss functions, including a Boundary Observant loss and a novel regularization term designed to enforce temporal consistency. Finally, we assess model robustness through controlled temporal distortions demonstrating that the learned representations rely predominantly on visual content rather than implicit. Overall, this work highlights the importance of structured temporal representations and tailored training objectives for accurate action progress estimation, and provides insights into design ing models that generalize across diverse domains and temporal conditions.| File | Dimensione | Formato | |
|---|---|---|---|
|
Mattesco_Antonio.pdf
accesso aperto
Dimensione
3.69 MB
Formato
Adobe PDF
|
3.69 MB | Adobe PDF | Visualizza/Apri |
The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License
https://hdl.handle.net/20.500.12608/110930