Back to search

University of Illinois Urbana-Champaign

Multimodal emotion recognition and speaker identification in financial conversations

Abstract

dc:description

This dissertation explores the use of multimodal, speaker-, and emotion-aware modeling for financial discourse, with a primary focus on Multimodal Emotion Recognition in Conversation and its downstream impact on key financial inference tasks. Motivated by the limitations of unimodal sentiment analysis and the strategic language regulation present in corporate communication, we introduce a large-scale, Large Language Model-based annotation framework for labeling emotion and emotion intensity in quarterly earnings call transcripts and audio recordings. We conduct a rigorous evaluation of annotation quality under varying prompt configurations and deterministic settings, uncovering key trade-offs between diversity and reliability. The resulting novel corpus, Multimodal Financial Emotion, serves as the foundation for developing MERSI, a multimodal, multi-task model that jointly predicts emotional state and speaker identity, leveraging contextual and acoustic cues. We show that this joint formulation improves performance over unimodal and context-agnostic baselines, particularly in capturing the nuanced structure of multi-label emotion recognition. Building on these insights, we evaluate MERSI in the scope of representation learning with respect to two downstream financial applications: Financial Restatement Prediction and Market Movement Prediction. Moreover, we introduce a Label Distribution Learning-based variant of our proposed model, which offers superior generalization on minority class predictions, highlighting its utility in capturing subtle emotional and narrative cues. Notably, the proposed models both exhibit stronger performance when predicting negative outcomes, which potentially reflects the greater emotional salience of adverse events. These findings further underscore the potential of multimodal, speaker-aware modeling as a scalable, generalizable framework not only for general emotion recognition but also in high-stakes financial decision-making and regulatory insight.

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Dissertation
Discipline thesis:degree_discipline
Informatics
Grantor
University of Illinois Urbana-Champaign
Year dc:date
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Kaikaus, Jamshed
Contributors dc:contributor
  • Brunner, Robert J.
  • Brunner, Robert J
  • Mendoza, Kimberly
  • Carrasco Kind, Matias
  • Zhu, Wei

Subjects

dc:subject × 16

Rights

dc:rights
Statement dc:rights
  • Copyright 2025 Jamshed Kaikaus
Language dc:language
en, eng

Identifiers

dc:identifier.*
Handle dc:identifier
https://hdl.handle.net/2142/129951

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Kaikaus, Jamshed. Multimodal emotion recognition and speaker identification in financial conversations. Dissertation thesis, University of Illinois Urbana-Champaign, 2025. https://hdl.handle.net/2142/129951