Back to search

University of Illinois Urbana-Champaign

Enhancing mid-level fusion with attention-based dual labeling for multimodal emotion recognition

Abstract

dc:description

Emotion recognition plays a pivotal role in human-computer interaction by enabling systems to interpret and adapt to users’ affective states. Traditional models typically rely on discrete categorical labels, which oversimplify the ambiguity and subjectivity inherent in human emotional expressions. This rigid labeling constrains model generalizability, especially in applications such as healthcare, social robotics, and virtual assistants, where nuanced understanding is essential. To address these limitations, this study introduces a probabilistic multimodal emotion recognition framework that combines attention-based mid-level fusion with dual-label learning. The model integrates audio embeddings from wav2vec2.0 and facial features from ResNet50-Face using a cross-modal attention mechanism that dynamically reweights modality contributions, enabling richer cross-modal interactions compared to early or late fusion strategies. Crucially, the framework incorporates a dual-label learning paradigm to jointly model self-reported (actor-intended) and observer-perceived emotions using soft probabilistic labels. A temperature-scaled softmax formulation captures uncertainty in emotion perception and improves interpretability by modeling distributions over emotion categories instead of committing to hard labels. Evaluation on the RAVDESS dataset, augmented through techniques such as pitch shifting, noise injection, and occlusion simulation, demonstrates the framework’s effectiveness. The model achieves 80.1% accuracy and 0.79 macro-F1 score, outperforming unimodal and simple fusion baselines. Additionally, it achieves a substantially lower KL divergence (0.51) and improved Expected Calibration Error (4.7%), indicating better alignment with human-perceived emotion distributions and improved prediction confidence calibration. These results highlight the promise of probabilistic multimodal learning for affective computing. By explicitly modeling uncertainty and leveraging both actor and observer perspectives, the proposed framework enables more robust, interpretable, and empathetic emotion-aware AI—paving the way for adaptive, emotionally intelligent systems in high-stakes domains.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Industrial Engineering
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Vasudeva, Sachit
Contributors dc:contributor
  • Kim, Inki

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Sachit Vasudeva
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/129791

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Vasudeva, Sachit. Enhancing mid-level fusion with attention-based dual labeling for multimodal emotion recognition. Thesis thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/129791