This work addresses continuous-valued valence and arousal estimation from physiological signals, evaluated on three corpora with different sensor coverage and annotation protocols, namely EmoWear, CASE, and RECOLA. Five physiological modalities, such as electrodermal activity, heart-rate variability, skin temperature, respiration, and blood-volume pulse, were processed into fractal, multifractal, and statistical handcrafted features, and combined with a raw-signal convolutional encoder. For RECOLA, the only corpus with paired speech recordings, a transformer-based speech branch pretrained for emotion recognition on RAVDESS was additionally fused into the model. The original contributions of this work are a multi-head self-attention mechanism that fuses the physiological and audio streams without assuming a fixed relationship between them, and a systematic evaluation of this model across configurations: physiological-only against audio-fused, and with or without domain-adversarial training. Evaluation was done under a leave-one-subject-out protocol, using both the concordance correlation coefficient (CCC) and F1 score. Across 78 subjects (EmoWear and CASE combined), the baseline configuration reaches a held-out test CCC of 0.137 for valence and 0.169 for arousal, and F1 of 0.574 and 0.596 respectively; enabling domain-adversarial training improves these to 0.146 and 0.172 (CCC) and 0.577 and 0.603 (F1). Trained and evaluated on each corpus in isolation, CASE reaches a substantially higher test CCC (0.283) than EmoWear (0.102), a difference traced to their differing annotation protocols.

This work addresses continuous-valued valence and arousal estimation from physiological signals, evaluated on three corpora with different sensor coverage and annotation protocols, namely EmoWear, CASE, and RECOLA. Five physiological modalities, such as electrodermal activity, heart-rate variability, skin temperature, respiration, and blood-volume pulse, were processed into fractal, multifractal, and statistical handcrafted features, and combined with a raw-signal convolutional encoder. For RECOLA, the only corpus with paired speech recordings, a transformer-based speech branch pretrained for emotion recognition on RAVDESS was additionally fused into the model. The original contributions of this work are a multi-head self-attention mechanism that fuses the physiological and audio streams without assuming a fixed relationship between them, and a systematic evaluation of this model across configurations: physiological-only against audio-fused, and with or without domain-adversarial training. Evaluation was done under a leave-one-subject-out protocol, using both the concordance correlation coefficient (CCC) and F1 score. Across 78 subjects (EmoWear and CASE combined), the baseline configuration reaches a held-out test CCC of 0.137 for valence and 0.169 for arousal, and F1 of 0.574 and 0.596 respectively; enabling domain-adversarial training improves these to 0.146 and 0.172 (CCC) and 0.577 and 0.603 (F1). Trained and evaluated on each corpus in isolation, CASE reaches a substantially higher test CCC (0.283) than EmoWear (0.102), a difference traced to their differing annotation protocols.

Multimodal Emotion Detection through Multichannel Signal Fusion

SHNAIDER, MARGARITA
2025/2026

Abstract

This work addresses continuous-valued valence and arousal estimation from physiological signals, evaluated on three corpora with different sensor coverage and annotation protocols, namely EmoWear, CASE, and RECOLA. Five physiological modalities, such as electrodermal activity, heart-rate variability, skin temperature, respiration, and blood-volume pulse, were processed into fractal, multifractal, and statistical handcrafted features, and combined with a raw-signal convolutional encoder. For RECOLA, the only corpus with paired speech recordings, a transformer-based speech branch pretrained for emotion recognition on RAVDESS was additionally fused into the model. The original contributions of this work are a multi-head self-attention mechanism that fuses the physiological and audio streams without assuming a fixed relationship between them, and a systematic evaluation of this model across configurations: physiological-only against audio-fused, and with or without domain-adversarial training. Evaluation was done under a leave-one-subject-out protocol, using both the concordance correlation coefficient (CCC) and F1 score. Across 78 subjects (EmoWear and CASE combined), the baseline configuration reaches a held-out test CCC of 0.137 for valence and 0.169 for arousal, and F1 of 0.574 and 0.596 respectively; enabling domain-adversarial training improves these to 0.146 and 0.172 (CCC) and 0.577 and 0.603 (F1). Trained and evaluated on each corpus in isolation, CASE reaches a substantially higher test CCC (0.283) than EmoWear (0.102), a difference traced to their differing annotation protocols.
2025
Multimodal Emotion Detection through Multichannel Signal Fusion
This work addresses continuous-valued valence and arousal estimation from physiological signals, evaluated on three corpora with different sensor coverage and annotation protocols, namely EmoWear, CASE, and RECOLA. Five physiological modalities, such as electrodermal activity, heart-rate variability, skin temperature, respiration, and blood-volume pulse, were processed into fractal, multifractal, and statistical handcrafted features, and combined with a raw-signal convolutional encoder. For RECOLA, the only corpus with paired speech recordings, a transformer-based speech branch pretrained for emotion recognition on RAVDESS was additionally fused into the model. The original contributions of this work are a multi-head self-attention mechanism that fuses the physiological and audio streams without assuming a fixed relationship between them, and a systematic evaluation of this model across configurations: physiological-only against audio-fused, and with or without domain-adversarial training. Evaluation was done under a leave-one-subject-out protocol, using both the concordance correlation coefficient (CCC) and F1 score. Across 78 subjects (EmoWear and CASE combined), the baseline configuration reaches a held-out test CCC of 0.137 for valence and 0.169 for arousal, and F1 of 0.574 and 0.596 respectively; enabling domain-adversarial training improves these to 0.146 and 0.172 (CCC) and 0.577 and 0.603 (F1). Trained and evaluated on each corpus in isolation, CASE reaches a substantially higher test CCC (0.283) than EmoWear (0.102), a difference traced to their differing annotation protocols.
Affective Computing
Fractal Features
Graph Neural Network
Emotion Detection
Multimodality
File in questo prodotto:
File Dimensione Formato  
Shnaider_Margarita.pdf

accesso aperto

Dimensione 16.53 MB
Formato Adobe PDF
16.53 MB Adobe PDF Visualizza/Apri

The text of this website © Università degli studi di Padova. Full Text are published under a non-exclusive license. Metadata are under a CC0 License

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/20.500.12608/113160