{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129791"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129791","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Enhancing mid-level fusion with attention-based dual labeling for multimodal emotion recognition","abstract":"Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-05-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;Closed Access&#x27;, the embargo will last until 2027-05-01","abstract_has_math":false,"creators":["Vasudeva, Sachit"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Industrial Engineering","degree_department":null,"school":null,"contributors":["Kim, Inki"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-05-09","date_published":"2025-05-09","updated_at":"2026-07-22T22:25:05Z","subjects":["Affective Computing","Cross-model Attention","Emotion Recognition"],"languages":["en","eng"],"rights":["Copyright 2025 Sachit Vasudeva"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129791","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kim, Inki"]},{"key":"dc:creator","label":"Author","values":["Vasudeva, Sachit"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-05-09","2025-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Industrial Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Affective Computing","Cross-model Attention","Emotion Recognition"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Sachit Vasudeva"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129791"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-05-01","The student, Sachit Vasudeva, accepted the attached license on 2025-05-09 at 14:56.","The student, Sachit Vasudeva, submitted this Thesis for approval on 2025-05-09 at 15:03.","This Thesis was approved for publication on 2025-05-09 at 16:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22291 on 2025-10-19 at 19:55:53","Emotion recognition plays a pivotal role in human-computer interaction by enabling systems to interpret and adapt to users’ affective states. Traditional models typically rely on discrete categorical labels, which oversimplify the ambiguity and subjectivity inherent in human emotional expressions. This rigid labeling constrains model generalizability, especially in applications such as healthcare, social robotics, and virtual assistants, where nuanced understanding is essential. To address these limitations, this study introduces a probabilistic multimodal emotion recognition framework that combines attention-based mid-level fusion with dual-label learning. The model integrates audio embeddings from wav2vec2.0 and facial features from ResNet50-Face using a cross-modal attention mechanism that dynamically reweights modality contributions, enabling richer cross-modal interactions compared to early or late fusion strategies. Crucially, the framework incorporates a dual-label learning paradigm to jointly model self-reported (actor-intended) and observer-perceived emotions using soft probabilistic labels. A temperature-scaled softmax formulation captures uncertainty in emotion perception and improves interpretability by modeling distributions over emotion categories instead of committing to hard labels. Evaluation on the RAVDESS dataset, augmented through techniques such as pitch shifting, noise injection, and occlusion simulation, demonstrates the framework’s effectiveness. The model achieves 80.1% accuracy and 0.79 macro-F1 score, outperforming unimodal and simple fusion baselines. Additionally, it achieves a substantially lower KL divergence (0.51) and improved Expected Calibration Error (4.7%), indicating better alignment with human-perceived emotion distributions and improved prediction confidence calibration. These results highlight the promise of probabilistic multimodal learning for affective computing. By explicitly modeling uncertainty and leveraging both actor and observer perspectives, the proposed framework enables more robust, interpretable, and empathetic emotion-aware AI—paving the way for adaptive, emotionally intelligent systems in high-stakes domains."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Enhancing mid-level fusion with attention-based dual labeling for multimodal emotion recognition"]}]}],"canonical_facts":{"dc:contributor":["Kim, Inki"],"dc:creator":["Vasudeva, Sachit"],"dc:date":["2025-05-09","2025-05"],"dc:description":["Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-05-01","The student, Sachit Vasudeva, accepted the attached license on 2025-05-09 at 14:56.","The student, Sachit Vasudeva, submitted this Thesis for approval on 2025-05-09 at 15:03.","This Thesis was approved for publication on 2025-05-09 at 16:36.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22291 on 2025-10-19 at 19:55:53","Emotion recognition plays a pivotal role in human-computer interaction by enabling systems to interpret and adapt to users’ affective states. Traditional models typically rely on discrete categorical labels, which oversimplify the ambiguity and subjectivity inherent in human emotional expressions. This rigid labeling constrains model generalizability, especially in applications such as healthcare, social robotics, and virtual assistants, where nuanced understanding is essential. To address these limitations, this study introduces a probabilistic multimodal emotion recognition framework that combines attention-based mid-level fusion with dual-label learning. The model integrates audio embeddings from wav2vec2.0 and facial features from ResNet50-Face using a cross-modal attention mechanism that dynamically reweights modality contributions, enabling richer cross-modal interactions compared to early or late fusion strategies. Crucially, the framework incorporates a dual-label learning paradigm to jointly model self-reported (actor-intended) and observer-perceived emotions using soft probabilistic labels. A temperature-scaled softmax formulation captures uncertainty in emotion perception and improves interpretability by modeling distributions over emotion categories instead of committing to hard labels. Evaluation on the RAVDESS dataset, augmented through techniques such as pitch shifting, noise injection, and occlusion simulation, demonstrates the framework’s effectiveness. The model achieves 80.1% accuracy and 0.79 macro-F1 score, outperforming unimodal and simple fusion baselines. Additionally, it achieves a substantially lower KL divergence (0.51) and improved Expected Calibration Error (4.7%), indicating better alignment with human-perceived emotion distributions and improved prediction confidence calibration. These results highlight the promise of probabilistic multimodal learning for affective computing. By explicitly modeling uncertainty and leveraging both actor and observer perspectives, the proposed framework enables more robust, interpretable, and empathetic emotion-aware AI—paving the way for adaptive, emotionally intelligent systems in high-stakes domains."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129791"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Sachit Vasudeva"],"dc:subject":["Affective Computing","Cross-model Attention","Emotion Recognition"],"dc:title":["Enhancing mid-level fusion with attention-based dual labeling for multimodal emotion recognition"],"dc:type":["text"],"thesis:degree_discipline":["Industrial Engineering"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:05Z"}