{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115460"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115460","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Guided text summarization with limited supervision","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Mao, Yuning"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Ji, Heng","Abdelzaher, Tarek","Yih, Wen-tau"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:54Z","subjects":["text summarization","natural language processing","text generation"],"languages":["en","eng"],"rights":["Copyright 2022 Yuning Mao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115460","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Ji, Heng","Abdelzaher, Tarek","Yih, Wen-tau"]},{"key":"dc:creator","label":"Author","values":["Mao, Yuning"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-15"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["text summarization","natural language processing","text generation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Yuning Mao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115460"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Yuning Mao, accepted the attached license on 2022-04-14 at 19:40.","The student, Yuning Mao, submitted this Dissertation for approval on 2022-04-14 at 19:48.","This Dissertation was approved for publication on 2022-04-15 at 14:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17681 on 2022-11-11 at 13:15:17","In the information age, people are surrounded by a massive amount of texts in daily life. It has become increasingly difficult to acquire desirable, salient, and non-redundant information from unstructured text data. Given the abundance of textual information, automatic text summarization rises as a critical means to provide people with salient and adequate information in a timely and manageable fashion. Text summarization has seen great progress in recent years, yet many issues remain or emerge, such as the lack of high-quality training data and model hallucination, due to the data-driven nature of mainstream summarization methods without necessary guidance. In light of the aforementioned background and challenges, my research aims to improve text summarization with more explicit guidance. In contrast to most prior studies on guided text summarization, I propose to conduct guided summarization with limited supervision – I design guided text summarization methods that either require limited labeled data or are totally unsupervised, which is critical for the development of text summarization as human annotations are typically expensive and time-consuming to obtain while the backbone of modern text summarizers (i.e., data-hungry neural models) require increasingly more training data. I have worked on guided text summarization from two aspects – the summarizer’s aspect and the user’s aspect, with various subjects as guidance (e.g., background documents and keyphrases). Summarizer-centric guidance directs summarizer with weak supervision, while user-centric guidance grants more control of the summarizer to a user with limited but valuable supervision from the user. From the summarizer’s perspective, mainstream summarization methods are typically data-driven without explicit guidance, which might be effective with abundant labeled data but fall short when the available training data is limited. I thus design methods that direct summarizer with weak supervision and improve summary quality without human annotation. The guidance received by a summarizer can be from internal sources, e.g., by leveraging content redundancy [1] and document structure [2]. The guidance can also be external, e.g., by making use of surrounding documents, which contain useful background information [3] or even provide direct supervision signals [4]. From the user’s perspective, mainstream summarizers are not customizable and generally offer the same output to different users. However, different users have various preferences and information needs in nature. I thus design methods that grant more control of the summarizer to the user so that a user can directly guide the summarization process and obtain tailored summaries suited to their own preferences, such as covering specific entities or events and generating summaries with different granularity [5, 6, 7]."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Guided text summarization with limited supervision"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Ji, Heng","Abdelzaher, Tarek","Yih, Wen-tau"],"dc:creator":["Mao, Yuning"],"dc:date":["2022-05","2022-04-15"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Yuning Mao, accepted the attached license on 2022-04-14 at 19:40.","The student, Yuning Mao, submitted this Dissertation for approval on 2022-04-14 at 19:48.","This Dissertation was approved for publication on 2022-04-15 at 14:02.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17681 on 2022-11-11 at 13:15:17","In the information age, people are surrounded by a massive amount of texts in daily life. It has become increasingly difficult to acquire desirable, salient, and non-redundant information from unstructured text data. Given the abundance of textual information, automatic text summarization rises as a critical means to provide people with salient and adequate information in a timely and manageable fashion. Text summarization has seen great progress in recent years, yet many issues remain or emerge, such as the lack of high-quality training data and model hallucination, due to the data-driven nature of mainstream summarization methods without necessary guidance. In light of the aforementioned background and challenges, my research aims to improve text summarization with more explicit guidance. In contrast to most prior studies on guided text summarization, I propose to conduct guided summarization with limited supervision – I design guided text summarization methods that either require limited labeled data or are totally unsupervised, which is critical for the development of text summarization as human annotations are typically expensive and time-consuming to obtain while the backbone of modern text summarizers (i.e., data-hungry neural models) require increasingly more training data. I have worked on guided text summarization from two aspects – the summarizer’s aspect and the user’s aspect, with various subjects as guidance (e.g., background documents and keyphrases). Summarizer-centric guidance directs summarizer with weak supervision, while user-centric guidance grants more control of the summarizer to a user with limited but valuable supervision from the user. From the summarizer’s perspective, mainstream summarization methods are typically data-driven without explicit guidance, which might be effective with abundant labeled data but fall short when the available training data is limited. I thus design methods that direct summarizer with weak supervision and improve summary quality without human annotation. The guidance received by a summarizer can be from internal sources, e.g., by leveraging content redundancy [1] and document structure [2]. The guidance can also be external, e.g., by making use of surrounding documents, which contain useful background information [3] or even provide direct supervision signals [4]. From the user’s perspective, mainstream summarizers are not customizable and generally offer the same output to different users. However, different users have various preferences and information needs in nature. I thus design methods that grant more control of the summarizer to the user so that a user can directly guide the summarization process and obtain tailored summaries suited to their own preferences, such as covering specific entities or events and generating summaries with different granularity [5, 6, 7]."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115460"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Yuning Mao"],"dc:subject":["text summarization","natural language processing","text generation"],"dc:title":["Guided text summarization with limited supervision"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:54Z"}