{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/109405"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/109405","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Model-based feature construction and text representation for social media analysis","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2021-03-04 without embargo terms","abstract_has_math":false,"creators":["Morales, Alex"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Zhai, ChengXiang","Han, Jiawei","Hockenmaier, Julia","Ungar, Lyle"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-03-05T21:38:12Z","date_published":"2021-03-05T21:38:12Z","updated_at":"2026-07-22T22:24:50Z","subjects":["machine learning","nlp","AI","features","feature construction","humor","feature development","truth discovery","hiv"],"languages":["en"],"rights":["Copyright 2020 Alex Morales"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/109405","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhai, ChengXiang","Han, Jiawei","Hockenmaier, Julia","Ungar, Lyle"]},{"key":"dc:creator","label":"Author","values":["Morales, Alex"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2021-03-05T21:38:12Z","2020-12-01","2020-12"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["machine learning","nlp","AI","features","feature construction","humor","feature development","truth discovery","hiv"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Alex Morales"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/109405"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","The student, Alex Morales, accepted the attached license on 2020-11-30 at 16:12.","Text representation is at the foundation of most text-based applications. Surface features are insufficient for many tasks and therefore constructing powerful discriminative features in a general way is an open challenge. Current approaches use deep neural networks to bypass feature construction. While deep learning can learn sophisticated representations from the text, it requires a lot of training data, which might not be readily available, and the derived features are not necessarily interpretable. In this work, we explore a novel paradigm, model-based feature construction (MBFC), that allows us to construct semantic features that can potentially improve many applications. In brief, MBFC uses human knowledge and expertise as well as big data to guide the design of models that enhance predictive modeling and support the data mining process by extracting useful knowledge, which in turn can be used as features for downstream prediction tasks. In this dissertation, we show how this paradigm can be applied to several tasks of social media analysis. We explore how MBFC can be used to solve the problem of target misalignment for prediction, where the output variable and the data may be at different levels of resolution and the goal is to construct features that can bridge this gap. The MBFC method allows us to use additional related data, e.g. associated context, to facilitate semantic analysis and feature construction. In this dissertation, we focus on a subset of problems in which social media data, in particular text data, can be leveraged to construct useful representations for prediction. We explore several kinds of user-generated content in social media data such as review data for useful review prediction, micro-blogging data for urgent health-based prediction tasks, and discussion forum data for expert prediction. First, we propose a background mixture model to capture incongruity features in text and use these features for humor detection in restaurant reviews. Second, we propose a source reliability feature representation method for trustworthy comment identification that incorporates user aspect expertise when modeling fine-grained reliabilities in an online discussion forum. And finally, we propose multi-view attribute features that adapt MBFC to handle the target misalignment problem for topic-based features and apply this to tweets in order to forecast new diagnosis rates for sexually transmitted infections.","The student, Alex Morales, submitted this Dissertation for approval on 2020-11-30 at 16:14.","This Dissertation was approved for publication on 2020-12-01 at 16:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15995 on 2021-03-04 at 15:35:33","Made available in DSpace on 2021-03-05T21:38:12Z (GMT). No. of bitstreams: 3 MORALES-DISSERTATION-2020.pdf: 5783300 bytes, checksum: 385af3b1ee31066bd6825c76471275e7 (MD5) LICENSE.txt: 4209 bytes, checksum: c0ca6fac168ee30dbfbb18848fcc975a (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: b597eba869baa584ff37b73a62f3a857 (MD5) Previous issue date: 2020-12-01"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Model-based feature construction and text representation for social media analysis"]}]}],"canonical_facts":{"dc:contributor":["Zhai, ChengXiang","Han, Jiawei","Hockenmaier, Julia","Ungar, Lyle"],"dc:creator":["Morales, Alex"],"dc:date":["2021-03-05T21:38:12Z","2020-12-01","2020-12"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2021-03-04 without embargo terms","The student, Alex Morales, accepted the attached license on 2020-11-30 at 16:12.","Text representation is at the foundation of most text-based applications. Surface features are insufficient for many tasks and therefore constructing powerful discriminative features in a general way is an open challenge. Current approaches use deep neural networks to bypass feature construction. While deep learning can learn sophisticated representations from the text, it requires a lot of training data, which might not be readily available, and the derived features are not necessarily interpretable. In this work, we explore a novel paradigm, model-based feature construction (MBFC), that allows us to construct semantic features that can potentially improve many applications. In brief, MBFC uses human knowledge and expertise as well as big data to guide the design of models that enhance predictive modeling and support the data mining process by extracting useful knowledge, which in turn can be used as features for downstream prediction tasks. In this dissertation, we show how this paradigm can be applied to several tasks of social media analysis. We explore how MBFC can be used to solve the problem of target misalignment for prediction, where the output variable and the data may be at different levels of resolution and the goal is to construct features that can bridge this gap. The MBFC method allows us to use additional related data, e.g. associated context, to facilitate semantic analysis and feature construction. In this dissertation, we focus on a subset of problems in which social media data, in particular text data, can be leveraged to construct useful representations for prediction. We explore several kinds of user-generated content in social media data such as review data for useful review prediction, micro-blogging data for urgent health-based prediction tasks, and discussion forum data for expert prediction. First, we propose a background mixture model to capture incongruity features in text and use these features for humor detection in restaurant reviews. Second, we propose a source reliability feature representation method for trustworthy comment identification that incorporates user aspect expertise when modeling fine-grained reliabilities in an online discussion forum. And finally, we propose multi-view attribute features that adapt MBFC to handle the target misalignment problem for topic-based features and apply this to tweets in order to forecast new diagnosis rates for sexually transmitted infections.","The student, Alex Morales, submitted this Dissertation for approval on 2020-11-30 at 16:14.","This Dissertation was approved for publication on 2020-12-01 at 16:12.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15995 on 2021-03-04 at 15:35:33","Made available in DSpace on 2021-03-05T21:38:12Z (GMT). No. of bitstreams: 3 MORALES-DISSERTATION-2020.pdf: 5783300 bytes, checksum: 385af3b1ee31066bd6825c76471275e7 (MD5) LICENSE.txt: 4209 bytes, checksum: c0ca6fac168ee30dbfbb18848fcc975a (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: b597eba869baa584ff37b73a62f3a857 (MD5) Previous issue date: 2020-12-01"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/109405"],"dc:language":["en"],"dc:rights":["Copyright 2020 Alex Morales"],"dc:subject":["machine learning","nlp","AI","features","feature construction","humor","feature development","truth discovery","hiv"],"dc:title":["Model-based feature construction and text representation for social media analysis"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:50Z"}