{"id":{"repo_id":"utc","oai_identifier":"oai:scholar.utc.edu:theses-2236"},"canonical_url":"https://search.dev.ndltd.org/etd/utc/oai:scholar.utc.edu:theses-2236","repository":{"repo_id":"utc","name":"University of Tennessee - Chattanooga","base_url":"https://scholar.utc.edu/do/oai/"},"display":{"title":"Personalized and adaptive therapeutic music generation from biosignals using knowledge-guided multimodal large language models","abstract":"Multimodal generative models are reshaping digital therapeutics by enabling real-time synthesis of personalized content aligned with a user’s physiological and affective state. However, existing systems remain fragmented across modalities and often lack a unified framework that can jointly represent biosignals, natural language intent, music, and video under long-context constraints. This dissertation presents an end-to-end multimodal large language model ecosystem for personalized therapeutic music generation that combines discrete tokenization, evidence-grounded reasoning, and stabilized preference alignment. At the representation layer, the dissertation develops a family of tokenizers that convert continuous biomedical and media signals into compact discrete sequences for transformer-based modeling. Harmonizer provides high-fidelity music tokenization for stable long-horizon conditioning and controllable generation. EEG-Harmonizer introduces neural tokenization with a Token-Transformer and Electrode-Aware Importance mechanism, achieving 99.97% classification accuracy while maintaining strong performance with only 50% of electrodes. Biomedical-Harmonizer extends the same encoder--quantizer--decoder principles to multi-lead ECG, producing morphology-preserving token sequences for physiological conditioning and safety-aware personalization. Video-Harmonizer, a quick-learning tokenizer framework that reduces learning time by 83.3%, extends the ecosystem to ultra-high-resolution visual data, and supports future EEG- and ECG-conditioned audiovisual therapy generation. At the reasoning layer, Qmusic-MLLM unifies text, EEG, ECG, and music token spaces within a single generative framework. Patient or session context is grounded through retrieval-augmented generation, and a chain-of-thought planner produces concise therapeutic conditioning plans before long-horizon music-token generation. To support adaptive personalization without expensive token-level reinforcement learning, the system also introduces a contextual-bandit prompt-bank mechanism that selects interpretable prompts using EEG-grounded reward signals. To improve reliability under preference learning, this dissertation further proposes Hallucination-Suppressed Preference Optimization, a reference-anchored alignment method that constrains updates relative to a frozen pretrained model while improving preference adherence. Together, these contributions establish a scalable foundation for EEG- and ECG-conditioned therapeutic generation by combining high-fidelity tokenization, grounded reasoning, adaptive personalization, and stabilized multimodal preference optimization within the Qmusic-MLLM ecosystem.","abstract_html":"Multimodal generative models are reshaping digital therapeutics by enabling real-time synthesis of personalized content aligned with a user’s physiological and affective state. However, existing systems remain fragmented across modalities and often lack a unified framework that can jointly represent biosignals, natural language intent, music, and video under long-context constraints. This dissertation presents an end-to-end multimodal large language model ecosystem for personalized therapeutic music generation that combines discrete tokenization, evidence-grounded reasoning, and stabilized preference alignment. At the representation layer, the dissertation develops a family of tokenizers that convert continuous biomedical and media signals into compact discrete sequences for transformer-based modeling. Harmonizer provides high-fidelity music tokenization for stable long-horizon conditioning and controllable generation. EEG-Harmonizer introduces neural tokenization with a Token-Transformer and Electrode-Aware Importance mechanism, achieving 99.97% classification accuracy while maintaining strong performance with only 50% of electrodes. Biomedical-Harmonizer extends the same encoder--quantizer--decoder principles to multi-lead ECG, producing morphology-preserving token sequences for physiological conditioning and safety-aware personalization. Video-Harmonizer, a quick-learning tokenizer framework that reduces learning time by 83.3%, extends the ecosystem to ultra-high-resolution visual data, and supports future EEG- and ECG-conditioned audiovisual therapy generation. At the reasoning layer, Qmusic-MLLM unifies text, EEG, ECG, and music token spaces within a single generative framework. Patient or session context is grounded through retrieval-augmented generation, and a chain-of-thought planner produces concise therapeutic conditioning plans before long-horizon music-token generation. To support adaptive personalization without expensive token-level reinforcement learning, the system also introduces a contextual-bandit prompt-bank mechanism that selects interpretable prompts using EEG-grounded reward signals. To improve reliability under preference learning, this dissertation further proposes Hallucination-Suppressed Preference Optimization, a reference-anchored alignment method that constrains updates relative to a frozen pretrained model while improving preference adherence. Together, these contributions establish a scalable foundation for EEG- and ECG-conditioned therapeutic generation by combining high-fidelity tokenization, grounded reasoning, adaptive personalization, and stabilized multimodal preference optimization within the Qmusic-MLLM ecosystem.","abstract_has_math":false,"creators":["Amiri, Amin"],"institution":"University of Tennessee at Chattanooga","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Liang, Yu","Wu, Dalei; Nasab, Ahad; Heath, Gregory","College of Engineering and Computer Science"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":null,"date_issued":"","date_published":null,"updated_at":"2026-07-24T05:47:28Z","subjects":["Artificial intelligence","Combined modality therapy"],"languages":["English","eng"],"rights":[],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://scholar.utc.edu/theses/1065","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Liang, Yu","Wu, Dalei; Nasab, Ahad; Heath, Gregory","College of Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Amiri, Amin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-05-01T07:00:00Z"]},{"key":"dc:publisher","label":"Institution","values":["University of Tennessee at Chattanooga","Chattanooga (Tenn.)"]},{"key":"dc:relation","label":"Dc Relation","values":["Masters Theses and Doctoral Dissertations"]},{"key":"dc:type","label":"Dc Type","values":["Doctoral dissertations","Text"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Artificial intelligence","Combined modality therapy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholar.utc.edu/theses/1065"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Dept. of Computer Science and Engineering","Ph. D.; A dissertation submitted to the faculty of the University of Tennessee at Chattanooga in partial fulfillment of the requirements of the degree of Doctor of Philosophy."]},{"key":"dc:description.abstract","label":"Abstract","values":["Multimodal generative models are reshaping digital therapeutics by enabling real-time synthesis of personalized content aligned with a user’s physiological and affective state. However, existing systems remain fragmented across modalities and often lack a unified framework that can jointly represent biosignals, natural language intent, music, and video under long-context constraints. This dissertation presents an end-to-end multimodal large language model ecosystem for personalized therapeutic music generation that combines discrete tokenization, evidence-grounded reasoning, and stabilized preference alignment. At the representation layer, the dissertation develops a family of tokenizers that convert continuous biomedical and media signals into compact discrete sequences for transformer-based modeling. Harmonizer provides high-fidelity music tokenization for stable long-horizon conditioning and controllable generation. EEG-Harmonizer introduces neural tokenization with a Token-Transformer and Electrode-Aware Importance mechanism, achieving 99.97% classification accuracy while maintaining strong performance with only 50% of electrodes. Biomedical-Harmonizer extends the same encoder--quantizer--decoder principles to multi-lead ECG, producing morphology-preserving token sequences for physiological conditioning and safety-aware personalization. Video-Harmonizer, a quick-learning tokenizer framework that reduces learning time by 83.3%, extends the ecosystem to ultra-high-resolution visual data, and supports future EEG- and ECG-conditioned audiovisual therapy generation. At the reasoning layer, Qmusic-MLLM unifies text, EEG, ECG, and music token spaces within a single generative framework. Patient or session context is grounded through retrieval-augmented generation, and a chain-of-thought planner produces concise therapeutic conditioning plans before long-horizon music-token generation. To support adaptive personalization without expensive token-level reinforcement learning, the system also introduces a contextual-bandit prompt-bank mechanism that selects interpretable prompts using EEG-grounded reward signals. To improve reliability under preference learning, this dissertation further proposes Hallucination-Suppressed Preference Optimization, a reference-anchored alignment method that constrains updates relative to a frozen pretrained model while improving preference adherence. Together, these contributions establish a scalable foundation for EEG- and ECG-conditioned therapeutic generation by combining high-fidelity tokenization, grounded reasoning, adaptive personalization, and stabilized multimodal preference optimization within the Qmusic-MLLM ecosystem."]},{"key":"dc:title","label":"Title","values":["Personalized and adaptive therapeutic music generation from biosignals using knowledge-guided multimodal large language models"]}]}],"canonical_facts":{"dc:contributor":["Liang, Yu","Wu, Dalei; Nasab, Ahad; Heath, Gregory","College of Engineering and Computer Science"],"dc:creator":["Amiri, Amin"],"dc:date":["2026-05-01T07:00:00Z"],"dc:description":["Dept. of Computer Science and Engineering","Ph. D.; A dissertation submitted to the faculty of the University of Tennessee at Chattanooga in partial fulfillment of the requirements of the degree of Doctor of Philosophy."],"dc:description.abstract":["Multimodal generative models are reshaping digital therapeutics by enabling real-time synthesis of personalized content aligned with a user’s physiological and affective state. However, existing systems remain fragmented across modalities and often lack a unified framework that can jointly represent biosignals, natural language intent, music, and video under long-context constraints. This dissertation presents an end-to-end multimodal large language model ecosystem for personalized therapeutic music generation that combines discrete tokenization, evidence-grounded reasoning, and stabilized preference alignment. At the representation layer, the dissertation develops a family of tokenizers that convert continuous biomedical and media signals into compact discrete sequences for transformer-based modeling. Harmonizer provides high-fidelity music tokenization for stable long-horizon conditioning and controllable generation. EEG-Harmonizer introduces neural tokenization with a Token-Transformer and Electrode-Aware Importance mechanism, achieving 99.97% classification accuracy while maintaining strong performance with only 50% of electrodes. Biomedical-Harmonizer extends the same encoder--quantizer--decoder principles to multi-lead ECG, producing morphology-preserving token sequences for physiological conditioning and safety-aware personalization. Video-Harmonizer, a quick-learning tokenizer framework that reduces learning time by 83.3%, extends the ecosystem to ultra-high-resolution visual data, and supports future EEG- and ECG-conditioned audiovisual therapy generation. At the reasoning layer, Qmusic-MLLM unifies text, EEG, ECG, and music token spaces within a single generative framework. Patient or session context is grounded through retrieval-augmented generation, and a chain-of-thought planner produces concise therapeutic conditioning plans before long-horizon music-token generation. To support adaptive personalization without expensive token-level reinforcement learning, the system also introduces a contextual-bandit prompt-bank mechanism that selects interpretable prompts using EEG-grounded reward signals. To improve reliability under preference learning, this dissertation further proposes Hallucination-Suppressed Preference Optimization, a reference-anchored alignment method that constrains updates relative to a frozen pretrained model while improving preference adherence. Together, these contributions establish a scalable foundation for EEG- and ECG-conditioned therapeutic generation by combining high-fidelity tokenization, grounded reasoning, adaptive personalization, and stabilized multimodal preference optimization within the Qmusic-MLLM ecosystem."],"dc:identifier":["https://scholar.utc.edu/theses/1065"],"dc:language":["English","eng"],"dc:publisher":["University of Tennessee at Chattanooga","Chattanooga (Tenn.)"],"dc:relation":["Masters Theses and Doctoral Dissertations"],"dc:rights":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Artificial intelligence","Combined modality therapy"],"dc:title":["Personalized and adaptive therapeutic music generation from biosignals using knowledge-guided multimodal large language models"],"dc:type":["Doctoral dissertations","Text"]},"updated_at":"2026-07-24T05:47:28Z"}