Back to results

University of Tennessee at Chattanooga

Personalized and adaptive therapeutic music generation from biosignals using knowledge-guided multimodal large language models

Abstract

dc:description.abstract

Multimodal generative models are reshaping digital therapeutics by enabling real-time synthesis of personalized content aligned with a user’s physiological and affective state. However, existing systems remain fragmented across modalities and often lack a unified framework that can jointly represent biosignals, natural language intent, music, and video under long-context constraints. This dissertation presents an end-to-end multimodal large language model ecosystem for personalized therapeutic music generation that combines discrete tokenization, evidence-grounded reasoning, and stabilized preference alignment. At the representation layer, the dissertation develops a family of tokenizers that convert continuous biomedical and media signals into compact discrete sequences for transformer-based modeling. Harmonizer provides high-fidelity music tokenization for stable long-horizon conditioning and controllable generation. EEG-Harmonizer introduces neural tokenization with a Token-Transformer and Electrode-Aware Importance mechanism, achieving 99.97% classification accuracy while maintaining strong performance with only 50% of electrodes. Biomedical-Harmonizer extends the same encoder--quantizer--decoder principles to multi-lead ECG, producing morphology-preserving token sequences for physiological conditioning and safety-aware personalization. Video-Harmonizer, a quick-learning tokenizer framework that reduces learning time by 83.3%, extends the ecosystem to ultra-high-resolution visual data, and supports future EEG- and ECG-conditioned audiovisual therapy generation. At the reasoning layer, Qmusic-MLLM unifies text, EEG, ECG, and music token spaces within a single generative framework. Patient or session context is grounded through retrieval-augmented generation, and a chain-of-thought planner produces concise therapeutic conditioning plans before long-horizon music-token generation. To support adaptive personalization without expensive token-level reinforcement learning, the system also introduces a contextual-bandit prompt-bank mechanism that selects interpretable prompts using EEG-grounded reward signals. To improve reliability under preference learning, this dissertation further proposes Hallucination-Suppressed Preference Optimization, a reference-anchored alignment method that constrains updates relative to a frozen pretrained model while improving preference adherence. Together, these contributions establish a scalable foundation for EEG- and ECG-conditioned therapeutic generation by combining high-fidelity tokenization, grounded reasoning, adaptive personalization, and stabilized multimodal preference optimization within the Qmusic-MLLM ecosystem.

Degree

thesis:*
Grantor dc:publisher
University of Tennessee at Chattanooga

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Amiri, Amin
Contributors dc:contributor
  • Liang, Yu
  • Wu, Dalei; Nasab, Ahad; Heath, Gregory
  • College of Engineering and Computer Science

Subjects

dc:subject × 2

Rights

dc:rights
Language dc:language
English, eng

Identifiers

dc:identifier.*
Repository record dc:identifier
https://scholar.utc.edu/theses/1065
OAI identifier oai:identifier
oai:scholar.utc.edu:theses-2236

Chain of custody

source
Harvested from
University of Tennessee - Chattanooga
Base URL
scholar.utc.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Amiri, Amin. Personalized and adaptive therapeutic music generation from biosignals using knowledge-guided multimodal large language models. University of Tennessee at Chattanooga, https://scholar.utc.edu/theses/1065