Online action detection in assembly environments poses significant challenges for human-robot collaboration, requiring models capable of processing streaming video data in real time while maintaining high detection accuracy over long temporal horizons. Existing approaches, predominantly based on Transformer architectures such as LSTR, rely on self-attention mechanisms whose quadratic computational complexity limits their scalability and suitability for low-latency deployment. This thesis presents MADAM (Multimodal Action Detection for Assembly with Mamba), a novel online action detection framework that replaces the Transformer-based temporal modeling module with a stack of Structured State Space Model (SSM) layers based on the Mamba architecture. By processing temporal sequences of video features with linear complexity, MADAM enables efficient streaming inference through the maintenance and selective exponential decay of recurrent hidden states, preserving long-term temporal context without reprocessing past observations. To further enhance detection robustness, MADAM integrates a multimodal fusion strategy combining RGB visual features with complementary motion cues, either optical flow or 3D skeleton sequences. Notably, the skeleton-based modality achieves detection performance comparable to optical flow while eliminating the substantial computational overhead associated with flow estimation, rendering the system more practical for real-time industrial deployment. MADAM is evaluated on two benchmarks: THUMOS-14, a standard dataset for action detection, and ATTACH, a challenging assembly-specific multi-label dataset featuring fine-grained, simultaneously occurring actions annotated per hand. On ATTACH, MADAM achieves a frame mAP of 37.28% and an exact accuracy of 38.02%, outperforming competitive baselines such as LSTR and MiniROAD. Using skeleton features MADAM achieves a frame mAP of 37.53%, matching the optical flow variant. On THUMOS, MADAM achieves 66.13% mAP compared to 69.50% for LSTR. The model, in fact, demonstrates stronger performance relative to the baseline in multi-label settings than in single-label scenarios. These results demonstrate that SSM-based architectures represent a strong and efficient alternative to Transformer models for real-time action detection in industrial assembly scenarios.

Online action detection in assembly environments poses significant challenges for human-robot collaboration, requiring models capable of processing streaming video data in real time while maintaining high detection accuracy over long temporal horizons. Existing approaches, predominantly based on Transformer architectures such as LSTR, rely on self-attention mechanisms whose quadratic computational complexity limits their scalability and suitability for low-latency deployment. This thesis presents MADAM (Multimodal Action Detection for Assembly with Mamba), a novel online action detection framework that replaces the Transformer-based temporal modeling module with a stack of Structured State Space Model (SSM) layers based on the Mamba architecture. By processing temporal sequences of video features with linear complexity, MADAM enables efficient streaming inference through the maintenance and selective exponential decay of recurrent hidden states, preserving long-term temporal context without reprocessing past observations. To further enhance detection robustness, MADAM integrates a multimodal fusion strategy combining RGB visual features with complementary motion cues, either optical flow or 3D skeleton sequences. Notably, the skeleton-based modality achieves detection performance comparable to optical flow while eliminating the substantial computational overhead associated with flow estimation, rendering the system more practical for real-time industrial deployment. MADAM is evaluated on two benchmarks: THUMOS-14, a standard dataset for action detection, and ATTACH, a challenging assembly-specific multi-label dataset featuring fine-grained, simultaneously occurring actions annotated per hand. On ATTACH, MADAM achieves a frame mAP of 37.28% and an exact accuracy of 38.02%, outperforming competitive baselines such as LSTR and MiniROAD. Using skeleton features MADAM achieves a frame mAP of 37.53%, matching the optical flow variant. On THUMOS, MADAM achieves 66.13% mAP compared to 69.50% for LSTR. The model, in fact, demonstrates stronger performance relative to the baseline in multi-label settings than in single-label scenarios. These results demonstrate that SSM-based architectures represent a strong and efficient alternative to Transformer models for real-time action detection in industrial assembly scenarios.

MADAM: A Multimodal Mamba-Based Approach for Online Action Detection in Assembly Scenarios

CINEL, GIOVANNI
2025/2026

Abstract

Online action detection in assembly environments poses significant challenges for human-robot collaboration, requiring models capable of processing streaming video data in real time while maintaining high detection accuracy over long temporal horizons. Existing approaches, predominantly based on Transformer architectures such as LSTR, rely on self-attention mechanisms whose quadratic computational complexity limits their scalability and suitability for low-latency deployment. This thesis presents MADAM (Multimodal Action Detection for Assembly with Mamba), a novel online action detection framework that replaces the Transformer-based temporal modeling module with a stack of Structured State Space Model (SSM) layers based on the Mamba architecture. By processing temporal sequences of video features with linear complexity, MADAM enables efficient streaming inference through the maintenance and selective exponential decay of recurrent hidden states, preserving long-term temporal context without reprocessing past observations. To further enhance detection robustness, MADAM integrates a multimodal fusion strategy combining RGB visual features with complementary motion cues, either optical flow or 3D skeleton sequences. Notably, the skeleton-based modality achieves detection performance comparable to optical flow while eliminating the substantial computational overhead associated with flow estimation, rendering the system more practical for real-time industrial deployment. MADAM is evaluated on two benchmarks: THUMOS-14, a standard dataset for action detection, and ATTACH, a challenging assembly-specific multi-label dataset featuring fine-grained, simultaneously occurring actions annotated per hand. On ATTACH, MADAM achieves a frame mAP of 37.28% and an exact accuracy of 38.02%, outperforming competitive baselines such as LSTR and MiniROAD. Using skeleton features MADAM achieves a frame mAP of 37.53%, matching the optical flow variant. On THUMOS, MADAM achieves 66.13% mAP compared to 69.50% for LSTR. The model, in fact, demonstrates stronger performance relative to the baseline in multi-label settings than in single-label scenarios. These results demonstrate that SSM-based architectures represent a strong and efficient alternative to Transformer models for real-time action detection in industrial assembly scenarios.
2025
MADAM: A Multimodal Mamba-Based Approach for Online Action Detection in Assembly Scenarios
Online action detection in assembly environments poses significant challenges for human-robot collaboration, requiring models capable of processing streaming video data in real time while maintaining high detection accuracy over long temporal horizons. Existing approaches, predominantly based on Transformer architectures such as LSTR, rely on self-attention mechanisms whose quadratic computational complexity limits their scalability and suitability for low-latency deployment. This thesis presents MADAM (Multimodal Action Detection for Assembly with Mamba), a novel online action detection framework that replaces the Transformer-based temporal modeling module with a stack of Structured State Space Model (SSM) layers based on the Mamba architecture. By processing temporal sequences of video features with linear complexity, MADAM enables efficient streaming inference through the maintenance and selective exponential decay of recurrent hidden states, preserving long-term temporal context without reprocessing past observations. To further enhance detection robustness, MADAM integrates a multimodal fusion strategy combining RGB visual features with complementary motion cues, either optical flow or 3D skeleton sequences. Notably, the skeleton-based modality achieves detection performance comparable to optical flow while eliminating the substantial computational overhead associated with flow estimation, rendering the system more practical for real-time industrial deployment. MADAM is evaluated on two benchmarks: THUMOS-14, a standard dataset for action detection, and ATTACH, a challenging assembly-specific multi-label dataset featuring fine-grained, simultaneously occurring actions annotated per hand. On ATTACH, MADAM achieves a frame mAP of 37.28% and an exact accuracy of 38.02%, outperforming competitive baselines such as LSTR and MiniROAD. Using skeleton features MADAM achieves a frame mAP of 37.53%, matching the optical flow variant. On THUMOS, MADAM achieves 66.13% mAP compared to 69.50% for LSTR. The model, in fact, demonstrates stronger performance relative to the baseline in multi-label settings than in single-label scenarios. These results demonstrate that SSM-based architectures represent a strong and efficient alternative to Transformer models for real-time action detection in industrial assembly scenarios.
Computer Vision
Mamba
OAD
Assembly Scenario
File in questo prodotto:
File Dimensione Formato  
Cinel_Giovanni.pdf

Accesso riservato

Dimensione 2.8 MB
Formato Adobe PDF
2.8 MB Adobe PDF

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/110132