{"id":{"repo_id":"uwo","oai_identifier":"oai:uwo.scholaris.ca:20.500.14721/31848"},"canonical_url":"https://search.dev.ndltd.org/etd/uwo/oai:uwo.scholaris.ca:20.500.14721/31848","repository":{"repo_id":"uwo","name":"Western University","base_url":"https://uwo.scholaris.ca/server/oai/request"},"display":{"title":"Statistical Learning Methods for Challenges arised from Self-Reported Data","abstract":"This thesis focuses on developing advanced clustering methods and analyzing data arised from chronic pain (CP) studies, with a particular emphasis on the unique challenges posed by self-reported (SR) data. Latent class analysis (LCA) is explored in the early stages of this work to cluster patients, and the clusters are compared to find features that are significantly different among clusters. While LCA is effective for categorical variables, it fails to address the mixed data types and subjective biases inherent in SR data. To overcome these limitations, we propose a novel distance metric tailored specifically for SR questionnaire data. This distance incorporates the correlation distance with other elementary distances for clustering data of mixed type, which outperforms existing metrics in handling mixed data when SR variables are present. Additionally, interpretable clustering techniques are utilized to generate simple, actionable rules that can be applied in clinical practice. To integrate the domain knowledge of CP experts into the clustering process, a semi- supervised clustering algorithm is introduced, allowing the distance metric to be adjusted using pairwise constraints provided by CP experts. We develop a two-step active learning query strat- egy to identify and query the most informative patient cases, enhancing query efficiency and minimizing the number of interactions required between experts and the algorithm. In addition to clustering, we analyze data arised from CP studies and explore predictive modeling. Canonical correlation analysis (CCA) is applied to investigate relationships among CP measurements, revealing important connections between pain characteristics and psycho- logical factors. Furthermore, multiple classification models are used to predict nociplastic pain, and the best cut of each predictor is investigated using the prediction model. Overall, we made significant contributions to the field of CP studies by introducing novel methods for clustering CP patients and analyzing complex data relationships. The proposed approaches emphasize clinical applicability, interpretability, and the integration of domain knowledge, offering practical solutions for real-world challenges in CP management. These advancements provide a foundation for further exploration of personalized treatment strategies and an improved understanding of chronic pain mechanisms.","abstract_html":"This thesis focuses on developing advanced clustering methods and analyzing data arised from chronic pain (CP) studies, with a particular emphasis on the unique challenges posed by self-reported (SR) data. Latent class analysis (LCA) is explored in the early stages of this work to cluster patients, and the clusters are compared to find features that are significantly different among clusters. While LCA is effective for categorical variables, it fails to address the mixed data types and subjective biases inherent in SR data. To overcome these limitations, we propose a novel distance metric tailored specifically for SR questionnaire data. This distance incorporates the correlation distance with other elementary distances for clustering data of mixed type, which outperforms existing metrics in handling mixed data when SR variables are present. Additionally, interpretable clustering techniques are utilized to generate simple, actionable rules that can be applied in clinical practice. To integrate the domain knowledge of CP experts into the clustering process, a semi- supervised clustering algorithm is introduced, allowing the distance metric to be adjusted using pairwise constraints provided by CP experts. We develop a two-step active learning query strat- egy to identify and query the most informative patient cases, enhancing query efficiency and minimizing the number of interactions required between experts and the algorithm. In addition to clustering, we analyze data arised from CP studies and explore predictive modeling. Canonical correlation analysis (CCA) is applied to investigate relationships among CP measurements, revealing important connections between pain characteristics and psycho- logical factors. Furthermore, multiple classification models are used to predict nociplastic pain, and the best cut of each predictor is investigated using the prediction model. Overall, we made significant contributions to the field of CP studies by introducing novel methods for clustering CP patients and analyzing complex data relationships. The proposed approaches emphasize clinical applicability, interpretability, and the integration of domain knowledge, offering practical solutions for real-world challenges in CP management. These advancements provide a foundation for further exploration of personalized treatment strategies and an improved understanding of chronic pain mechanisms.","abstract_has_math":false,"creators":["Deng, Gansen"],"institution":"The University of Western Ontario","degree_name":"Ph D","degree_level":null,"degree_discipline":"Statistics and Actuarial Sciences","degree_department":null,"school":null,"contributors":[],"advisors":["He, Wenqing","Kumbhare, Dinesh"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-04-09","date_published":"2025-04-09","updated_at":"2026-07-27T21:56:14Z","subjects":["Chronic pain","self-reported data","interpretable clustering","semi-supervised clustering","active learning"],"languages":["en_ca"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/20.500.14721/31848","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["He, Wenqing","Kumbhare, Dinesh"]},{"key":"dc:creator","label":"Author","values":["Deng, Gansen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-07-10T19:23:22Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-07-10T19:23:22Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-04-09"]},{"key":"dc:publisher","label":"Institution","values":["The University of Western Ontario"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics and Actuarial Sciences"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph D"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Chronic pain","self-reported data","interpretable clustering","semi-supervised clustering","active learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_ca"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/20.500.14721/31848"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The thesis cover page in the PDF document includes references to Western University’s previous institutional repository platform, known as Scholarship@Western, and links to that platform (beginning with ir.lib.uwo.ca). In citing or referring to this thesis, use the DOI or handle from this page instead. Sample citation: Author name, \"Thesis title.\" (Year). Western University Open Repository. https://doi.org/10.71858/123456."]},{"key":"dc:description.abstract","label":"Abstract","values":["This thesis focuses on developing advanced clustering methods and analyzing data arised from chronic pain (CP) studies, with a particular emphasis on the unique challenges posed by self-reported (SR) data. Latent class analysis (LCA) is explored in the early stages of this work to cluster patients, and the clusters are compared to find features that are significantly different among clusters. While LCA is effective for categorical variables, it fails to address the mixed data types and subjective biases inherent in SR data. To overcome these limitations, we propose a novel distance metric tailored specifically for SR questionnaire data. This distance incorporates the correlation distance with other elementary distances for clustering data of mixed type, which outperforms existing metrics in handling mixed data when SR variables are present. Additionally, interpretable clustering techniques are utilized to generate simple, actionable rules that can be applied in clinical practice. To integrate the domain knowledge of CP experts into the clustering process, a semi- supervised clustering algorithm is introduced, allowing the distance metric to be adjusted using pairwise constraints provided by CP experts. We develop a two-step active learning query strat- egy to identify and query the most informative patient cases, enhancing query efficiency and minimizing the number of interactions required between experts and the algorithm. In addition to clustering, we analyze data arised from CP studies and explore predictive modeling. Canonical correlation analysis (CCA) is applied to investigate relationships among CP measurements, revealing important connections between pain characteristics and psycho- logical factors. Furthermore, multiple classification models are used to predict nociplastic pain, and the best cut of each predictor is investigated using the prediction model. Overall, we made significant contributions to the field of CP studies by introducing novel methods for clustering CP patients and analyzing complex data relationships. The proposed approaches emphasize clinical applicability, interpretability, and the integration of domain knowledge, offering practical solutions for real-world challenges in CP management. These advancements provide a foundation for further exploration of personalized treatment strategies and an improved understanding of chronic pain mechanisms."]},{"key":"dc:title","label":"Title","values":["Statistical Learning Methods for Challenges arised from Self-Reported Data"]}]}],"canonical_facts":{"dc:contributor.advisor":["He, Wenqing","Kumbhare, Dinesh"],"dc:creator":["Deng, Gansen"],"dc:date.accessioned":["2025-07-10T19:23:22Z"],"dc:date.available":["2025-07-10T19:23:22Z"],"dc:date.issued":["2025-04-09"],"dc:description":["The thesis cover page in the PDF document includes references to Western University’s previous institutional repository platform, known as Scholarship@Western, and links to that platform (beginning with ir.lib.uwo.ca). In citing or referring to this thesis, use the DOI or handle from this page instead. Sample citation: Author name, \"Thesis title.\" (Year). Western University Open Repository. https://doi.org/10.71858/123456."],"dc:description.abstract":["This thesis focuses on developing advanced clustering methods and analyzing data arised from chronic pain (CP) studies, with a particular emphasis on the unique challenges posed by self-reported (SR) data. Latent class analysis (LCA) is explored in the early stages of this work to cluster patients, and the clusters are compared to find features that are significantly different among clusters. While LCA is effective for categorical variables, it fails to address the mixed data types and subjective biases inherent in SR data. To overcome these limitations, we propose a novel distance metric tailored specifically for SR questionnaire data. This distance incorporates the correlation distance with other elementary distances for clustering data of mixed type, which outperforms existing metrics in handling mixed data when SR variables are present. Additionally, interpretable clustering techniques are utilized to generate simple, actionable rules that can be applied in clinical practice. To integrate the domain knowledge of CP experts into the clustering process, a semi- supervised clustering algorithm is introduced, allowing the distance metric to be adjusted using pairwise constraints provided by CP experts. We develop a two-step active learning query strat- egy to identify and query the most informative patient cases, enhancing query efficiency and minimizing the number of interactions required between experts and the algorithm. In addition to clustering, we analyze data arised from CP studies and explore predictive modeling. Canonical correlation analysis (CCA) is applied to investigate relationships among CP measurements, revealing important connections between pain characteristics and psycho- logical factors. Furthermore, multiple classification models are used to predict nociplastic pain, and the best cut of each predictor is investigated using the prediction model. Overall, we made significant contributions to the field of CP studies by introducing novel methods for clustering CP patients and analyzing complex data relationships. The proposed approaches emphasize clinical applicability, interpretability, and the integration of domain knowledge, offering practical solutions for real-world challenges in CP management. These advancements provide a foundation for further exploration of personalized treatment strategies and an improved understanding of chronic pain mechanisms."],"dc:identifier.uri":["https://hdl.handle.net/20.500.14721/31848"],"dc:language.iso":["en_ca"],"dc:publisher":["The University of Western Ontario"],"dc:subject":["Chronic pain","self-reported data","interpretable clustering","semi-supervised clustering","active learning"],"dc:title":["Statistical Learning Methods for Challenges arised from Self-Reported Data"],"dc:type":["thesis"],"thesis:degree_discipline":["Statistics and Actuarial Sciences"],"thesis:degree_name":["Ph D"]},"updated_at":"2026-07-27T21:56:14Z"}