University of Tennessee at Chattanooga
Personalized and adaptive therapeutic music generation from biosignals using knowledge-guided multimodal large language models
Abstract
dc:description.abstractMultimodal generative models are reshaping digital therapeutics by enabling real-time synthesis of personalized content aligned with a user’s physiological and affective state. However, existing systems remain fragmented across modalities and often lack a unified framework that can jointly represent biosignals, natural language intent, music, and video under long-context constraints. This dissertation presents an end-to-end multimodal large language model ecosystem for personalized therapeutic music generation that combines discrete tokenization, evidence-grounded reasoning, and stabilized preference alignment. At the representation layer, the dissertation develops a family of tokenizers that convert continuous biomedical and media signals into compact discrete sequences for transformer-based modeling. Harmonizer provides high-fidelity music tokenization for stable long-horizon conditioning and controllable generation. EEG-Harmonizer introduces neural tokenization with a Token-Transformer and Electrode-Aware Importance mechanism, achieving 99.97% classification accuracy while maintaining strong performance with only 50% of electrodes. Biomedical-Harmonizer extends the same encoder--quantizer--decoder principles to multi-lead ECG, producing morphology-preserving token sequences for physiological conditioning and safety-aware personalization. Video-Harmonizer, a quick-learning tokenizer framework that reduces learning time by 83.3%, extends the ecosystem to ultra-high-resolution visual data, and supports future EEG- and ECG-conditioned audiovisual therapy generation. At the reasoning layer, Qmusic-MLLM unifies text, EEG, ECG, and music token spaces within a single generative framework. Patient or session context is grounded through retrieval-augmented generation, and a chain-of-thought planner produces concise therapeutic conditioning plans before long-horizon music-token generation. To support adaptive personalization without expensive token-level reinforcement learning, the system also introduces a contextual-bandit prompt-bank mechanism that selects interpretable prompts using EEG-grounded reward signals. To improve reliability under preference learning, this dissertation further proposes Hallucination-Suppressed Preference Optimization, a reference-anchored alignment method that constrains updates relative to a frozen pretrained model while improving preference adherence. Together, these contributions establish a scalable foundation for EEG- and ECG-conditioned therapeutic generation by combining high-fidelity tokenization, grounded reasoning, adaptive personalization, and stabilized multimodal preference optimization within the Qmusic-MLLM ecosystem.
Degree
thesis:*- Grantor dc:publisher
- University of Tennessee at Chattanooga
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Amiri, Amin
- Contributors dc:contributor
-
- Liang, Yu
- Wu, Dalei; Nasab, Ahad; Heath, Gregory
- College of Engineering and Computer Science
Subjects
dc:subject × 2Rights
dc:rights- Language dc:language
- English, eng
Identifiers
dc:identifier.*- Repository record dc:identifier
- https://scholar.utc.edu/theses/1065
- OAI identifier oai:identifier
- oai:scholar.utc.edu:theses-2236