{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/129951"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/129951","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Multimodal emotion recognition and speaker identification in financial conversations","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2025-10-20 without embargo terms","abstract_has_math":false,"creators":["Kaikaus, Jamshed"],"institution":"University of Illinois Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Informatics","degree_department":null,"school":null,"contributors":["Brunner, Robert J.","Brunner, Robert J","Mendoza, Kimberly","Carrasco Kind, Matias","Zhu, Wei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-07-17","date_published":"2025-07-17","updated_at":"2026-07-22T22:25:06Z","subjects":["Emotion Recognition","Speaker Identification","Multimodal Learning","Multi-task Learning","Machine Learning","Deep Learning","Affective Computing","Representation Learning","Large Language Models","Natural Language Processing","Speech Processing","Financial Artificial Intelligence","Financial Communication","Earnings Calls","Data Annotation","Dataset Curation"],"languages":["en","eng"],"rights":["Copyright 2025 Jamshed Kaikaus"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/129951","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Brunner, Robert J.","Brunner, Robert J","Mendoza, Kimberly","Carrasco Kind, Matias","Zhu, Wei"]},{"key":"dc:creator","label":"Author","values":["Kaikaus, Jamshed"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-07-17","2025-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Informatics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Emotion Recognition","Speaker Identification","Multimodal Learning","Multi-task Learning","Machine Learning","Deep Learning","Affective Computing","Representation Learning","Large Language Models","Natural Language Processing","Speech Processing","Financial Artificial Intelligence","Financial Communication","Earnings Calls","Data Annotation","Dataset Curation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Jamshed Kaikaus"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/129951"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Jamshed Kaikaus, accepted the attached license on 2025-07-16 at 15:29.","The student, Jamshed Kaikaus, submitted this Dissertation for approval on 2025-07-16 at 16:01.","This Dissertation was approved for publication on 2025-07-17 at 14:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22609 on 2025-10-20 at 20:15:22","This dissertation explores the use of multimodal, speaker-, and emotion-aware modeling for financial discourse, with a primary focus on Multimodal Emotion Recognition in Conversation and its downstream impact on key financial inference tasks. Motivated by the limitations of unimodal sentiment analysis and the strategic language regulation present in corporate communication, we introduce a large-scale, Large Language Model-based annotation framework for labeling emotion and emotion intensity in quarterly earnings call transcripts and audio recordings. We conduct a rigorous evaluation of annotation quality under varying prompt configurations and deterministic settings, uncovering key trade-offs between diversity and reliability. The resulting novel corpus, Multimodal Financial Emotion, serves as the foundation for developing MERSI, a multimodal, multi-task model that jointly predicts emotional state and speaker identity, leveraging contextual and acoustic cues. We show that this joint formulation improves performance over unimodal and context-agnostic baselines, particularly in capturing the nuanced structure of multi-label emotion recognition. Building on these insights, we evaluate MERSI in the scope of representation learning with respect to two downstream financial applications: Financial Restatement Prediction and Market Movement Prediction. Moreover, we introduce a Label Distribution Learning-based variant of our proposed model, which offers superior generalization on minority class predictions, highlighting its utility in capturing subtle emotional and narrative cues. Notably, the proposed models both exhibit stronger performance when predicting negative outcomes, which potentially reflects the greater emotional salience of adverse events. These findings further underscore the potential of multimodal, speaker-aware modeling as a scalable, generalizable framework not only for general emotion recognition but also in high-stakes financial decision-making and regulatory insight."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Multimodal emotion recognition and speaker identification in financial conversations"]}]}],"canonical_facts":{"dc:contributor":["Brunner, Robert J.","Brunner, Robert J","Mendoza, Kimberly","Carrasco Kind, Matias","Zhu, Wei"],"dc:creator":["Kaikaus, Jamshed"],"dc:date":["2025-07-17","2025-08"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2025-10-20 without embargo terms","The student, Jamshed Kaikaus, accepted the attached license on 2025-07-16 at 15:29.","The student, Jamshed Kaikaus, submitted this Dissertation for approval on 2025-07-16 at 16:01.","This Dissertation was approved for publication on 2025-07-17 at 14:07.","DSpace SAF Submission Ingestion Package generated from Vireo submission #22609 on 2025-10-20 at 20:15:22","This dissertation explores the use of multimodal, speaker-, and emotion-aware modeling for financial discourse, with a primary focus on Multimodal Emotion Recognition in Conversation and its downstream impact on key financial inference tasks. Motivated by the limitations of unimodal sentiment analysis and the strategic language regulation present in corporate communication, we introduce a large-scale, Large Language Model-based annotation framework for labeling emotion and emotion intensity in quarterly earnings call transcripts and audio recordings. We conduct a rigorous evaluation of annotation quality under varying prompt configurations and deterministic settings, uncovering key trade-offs between diversity and reliability. The resulting novel corpus, Multimodal Financial Emotion, serves as the foundation for developing MERSI, a multimodal, multi-task model that jointly predicts emotional state and speaker identity, leveraging contextual and acoustic cues. We show that this joint formulation improves performance over unimodal and context-agnostic baselines, particularly in capturing the nuanced structure of multi-label emotion recognition. Building on these insights, we evaluate MERSI in the scope of representation learning with respect to two downstream financial applications: Financial Restatement Prediction and Market Movement Prediction. Moreover, we introduce a Label Distribution Learning-based variant of our proposed model, which offers superior generalization on minority class predictions, highlighting its utility in capturing subtle emotional and narrative cues. Notably, the proposed models both exhibit stronger performance when predicting negative outcomes, which potentially reflects the greater emotional salience of adverse events. These findings further underscore the potential of multimodal, speaker-aware modeling as a scalable, generalizable framework not only for general emotion recognition but also in high-stakes financial decision-making and regulatory insight."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/129951"],"dc:language":["en","eng"],"dc:rights":["Copyright 2025 Jamshed Kaikaus"],"dc:subject":["Emotion Recognition","Speaker Identification","Multimodal Learning","Multi-task Learning","Machine Learning","Deep Learning","Affective Computing","Representation Learning","Large Language Models","Natural Language Processing","Speech Processing","Financial Artificial Intelligence","Financial Communication","Earnings Calls","Data Annotation","Dataset Curation"],"dc:title":["Multimodal emotion recognition and speaker identification in financial conversations"],"dc:type":["text"],"thesis:degree_discipline":["Informatics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:06Z"}