{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/42485"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/42485","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Clustering and comparing information extracted from personal health messages","abstract":"The development of Web 2.0 techniques has led to the prosperity of online communities, which spread to various domains and areas in our daily life. When it comes to the medicine and healthcare domain, a series of good online services such as Yahoo! Groups,WebMD and Med- Help, offer patients and physicians a good platform to discuss health problems, e.g., diseases and drugs, diagnoses and treatments, which also provide a large volume of data for researchers to analyze and explore. However, some nature of the personal messages, e.g., unclean, unstructured and isolated from clinical practice, hinders users’ effective digestion of information in the front end and challenges the data analysis in the back end. In such a scenario, the objective of my thesis is to apply the advanced data mining, information retrieval and natural language processing techniques to effectively analyze and re-organize the rich source of personal health messages from online medical communities, in order to satisfy patients’ information need and support physicians’ clinical practice. Specially, in the first part of the dissertation, I introduce an SVM-based multi-class classification method which utilizes term-appearance, lexical and semantic features to effectively classify health messages sampled from our unique dataset of Yahoo! Health Groups into three categories: News, User Comments and Spam; in the second part, I depict a comprehensive system with an extensive evaluation framework to organize and cluster patient outcomes utilizing topic model, which groups large collections of personal comments into a series of topics, guided by expert comments; in the third part of the dissertation, I address a novel and promising topic: Comparative Effectiveness Research (CER) hypothesis prediction, by presenting a study which evaluates patients’ opinions on different treatments by machine enabled sentiment analysis or human analysts utilizing our MedHelp dataset. By suggesting three different methods to compare such opinions, reliable conclusions about the patients’ preference on different treatments can be drawn consistently, which imply the effectiveness of the treatments. Furthermore, the study is also extended to demographic analysis to explore the preference in specific group of people, representing population cohorts.","abstract_html":"The development of Web 2.0 techniques has led to the prosperity of online communities, which spread to various domains and areas in our daily life. When it comes to the medicine and healthcare domain, a series of good online services such as Yahoo! Groups,WebMD and Med- Help, offer patients and physicians a good platform to discuss health problems, e.g., diseases and drugs, diagnoses and treatments, which also provide a large volume of data for researchers to analyze and explore. However, some nature of the personal messages, e.g., unclean, unstructured and isolated from clinical practice, hinders users’ effective digestion of information in the front end and challenges the data analysis in the back end. In such a scenario, the objective of my thesis is to apply the advanced data mining, information retrieval and natural language processing techniques to effectively analyze and re-organize the rich source of personal health messages from online medical communities, in order to satisfy patients’ information need and support physicians’ clinical practice. Specially, in the first part of the dissertation, I introduce an SVM-based multi-class classification method which utilizes term-appearance, lexical and semantic features to effectively classify health messages sampled from our unique dataset of Yahoo! Health Groups into three categories: News, User Comments and Spam; in the second part, I depict a comprehensive system with an extensive evaluation framework to organize and cluster patient outcomes utilizing topic model, which groups large collections of personal comments into a series of topics, guided by expert comments; in the third part of the dissertation, I address a novel and promising topic: Comparative Effectiveness Research (CER) hypothesis prediction, by presenting a study which evaluates patients’ opinions on different treatments by machine enabled sentiment analysis or human analysts utilizing our MedHelp dataset. By suggesting three different methods to compare such opinions, reliable conclusions about the patients’ preference on different treatments can be drawn consistently, which imply the effectiveness of the treatments. Furthermore, the study is also extended to demographic analysis to explore the preference in specific group of people, representing population cohorts.","abstract_has_math":false,"creators":["Jiang, Yunliang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Schatz, Bruce R.","Han, Jiawei","Zhai, ChengXiang","Mei, Qiaozhu"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-02-03T19:47:26Z","date_published":"2013-02-03T19:47:26Z","updated_at":"2026-07-22T22:25:33Z","subjects":["Healthcare","Personal Messages","Classification","Clustering","Comparison"],"languages":["en"],"rights":["Copyright 2012 Yunliang Jiang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/42485","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Schatz, Bruce R.","Han, Jiawei","Zhai, ChengXiang","Mei, Qiaozhu"]},{"key":"dc:creator","label":"Author","values":["Jiang, Yunliang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-02-03T19:47:26Z","2015-02-03T11:00:59Z","2012-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Healthcare","Personal Messages","Classification","Clustering","Comparison"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Yunliang Jiang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/42485"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The development of Web 2.0 techniques has led to the prosperity of online communities, which spread to various domains and areas in our daily life. When it comes to the medicine and healthcare domain, a series of good online services such as Yahoo! Groups,WebMD and Med- Help, offer patients and physicians a good platform to discuss health problems, e.g., diseases and drugs, diagnoses and treatments, which also provide a large volume of data for researchers to analyze and explore. However, some nature of the personal messages, e.g., unclean, unstructured and isolated from clinical practice, hinders users’ effective digestion of information in the front end and challenges the data analysis in the back end. In such a scenario, the objective of my thesis is to apply the advanced data mining, information retrieval and natural language processing techniques to effectively analyze and re-organize the rich source of personal health messages from online medical communities, in order to satisfy patients’ information need and support physicians’ clinical practice. Specially, in the first part of the dissertation, I introduce an SVM-based multi-class classification method which utilizes term-appearance, lexical and semantic features to effectively classify health messages sampled from our unique dataset of Yahoo! Health Groups into three categories: News, User Comments and Spam; in the second part, I depict a comprehensive system with an extensive evaluation framework to organize and cluster patient outcomes utilizing topic model, which groups large collections of personal comments into a series of topics, guided by expert comments; in the third part of the dissertation, I address a novel and promising topic: Comparative Effectiveness Research (CER) hypothesis prediction, by presenting a study which evaluates patients’ opinions on different treatments by machine enabled sentiment analysis or human analysts utilizing our MedHelp dataset. By suggesting three different methods to compare such opinions, reliable conclusions about the patients’ preference on different treatments can be drawn consistently, which imply the effectiveness of the treatments. Furthermore, the study is also extended to demographic analysis to explore the preference in specific group of people, representing population cohorts.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-11-26T19:23:00Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Jiang_Yunliang.pdf: 738179 bytes, checksum: 81d4793eb8d36e306d6f046549ad3fcd (MD5)","Made available in DSpace on 2013-02-03T19:47:26Z (GMT). No. of bitstreams: 2 Yunliang_Jiang.pdf: 738179 bytes, checksum: 81d4793eb8d36e306d6f046549ad3fcd (MD5) license.txt: 4062 bytes, checksum: 85fdd2d8cfdf776f7c151963ea518d07 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (srobbins@illinois.edu) on 2013-02-03T19:48:03Z Item is restricted until 2015-02-03T19:47:48Z","Restriction data tranferred 2014-07-01T11:36:05-05:00 Original Data Group with Access Administrator Release Date: 2015-02-03 13:47:48 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 42433 on 2015-02-03T11:00:59Z."]},{"key":"dc:title","label":"Title","values":["Clustering and comparing information extracted from personal health messages"]}]}],"canonical_facts":{"dc:contributor":["Schatz, Bruce R.","Han, Jiawei","Zhai, ChengXiang","Mei, Qiaozhu"],"dc:creator":["Jiang, Yunliang"],"dc:date":["2013-02-03T19:47:26Z","2015-02-03T11:00:59Z","2012-12"],"dc:description":["The development of Web 2.0 techniques has led to the prosperity of online communities, which spread to various domains and areas in our daily life. When it comes to the medicine and healthcare domain, a series of good online services such as Yahoo! Groups,WebMD and Med- Help, offer patients and physicians a good platform to discuss health problems, e.g., diseases and drugs, diagnoses and treatments, which also provide a large volume of data for researchers to analyze and explore. However, some nature of the personal messages, e.g., unclean, unstructured and isolated from clinical practice, hinders users’ effective digestion of information in the front end and challenges the data analysis in the back end. In such a scenario, the objective of my thesis is to apply the advanced data mining, information retrieval and natural language processing techniques to effectively analyze and re-organize the rich source of personal health messages from online medical communities, in order to satisfy patients’ information need and support physicians’ clinical practice. Specially, in the first part of the dissertation, I introduce an SVM-based multi-class classification method which utilizes term-appearance, lexical and semantic features to effectively classify health messages sampled from our unique dataset of Yahoo! Health Groups into three categories: News, User Comments and Spam; in the second part, I depict a comprehensive system with an extensive evaluation framework to organize and cluster patient outcomes utilizing topic model, which groups large collections of personal comments into a series of topics, guided by expert comments; in the third part of the dissertation, I address a novel and promising topic: Comparative Effectiveness Research (CER) hypothesis prediction, by presenting a study which evaluates patients’ opinions on different treatments by machine enabled sentiment analysis or human analysts utilizing our MedHelp dataset. By suggesting three different methods to compare such opinions, reliable conclusions about the patients’ preference on different treatments can be drawn consistently, which imply the effectiveness of the treatments. Furthermore, the study is also extended to demographic analysis to explore the preference in specific group of people, representing population cohorts.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-11-26T19:23:00Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Jiang_Yunliang.pdf: 738179 bytes, checksum: 81d4793eb8d36e306d6f046549ad3fcd (MD5)","Made available in DSpace on 2013-02-03T19:47:26Z (GMT). No. of bitstreams: 2 Yunliang_Jiang.pdf: 738179 bytes, checksum: 81d4793eb8d36e306d6f046549ad3fcd (MD5) license.txt: 4062 bytes, checksum: 85fdd2d8cfdf776f7c151963ea518d07 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by Seth Robbins (srobbins@illinois.edu) on 2013-02-03T19:48:03Z Item is restricted until 2015-02-03T19:47:48Z","Restriction data tranferred 2014-07-01T11:36:05-05:00 Original Data Group with Access Administrator Release Date: 2015-02-03 13:47:48 UTC Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 42433 on 2015-02-03T11:00:59Z."],"dc:identifier":["http://hdl.handle.net/2142/42485"],"dc:language":["en"],"dc:rights":["Copyright 2012 Yunliang Jiang"],"dc:subject":["Healthcare","Personal Messages","Classification","Clustering","Comparison"],"dc:title":["Clustering and comparing information extracted from personal health messages"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:33Z"}