{"id":{"repo_id":"uic","oai_identifier":"oai:figshare.com:article/32993957"},"canonical_url":"https://search.dev.ndltd.org/etd/uic/oai:figshare.com:article/32993957","repository":{"repo_id":"uic","name":"University of Illinois - Chicago","base_url":"https://api.figshare.com/v2/oai"},"display":{"title":"Towards Label-Efficient Learning: Exploiting Limited Labels and Abundant Unlabeled Data","abstract":"Efficiently exploiting limited labels alongside abundant unlabeled data is critical for advancing the capabilities of language models and their downstream systems. This thesis first presents novel methodologies for label-efficient learning, introducing three frameworks: (1) JointMatch, which employs collaborative pseudo-labeling to leverage limited labeled examples together with extensive unlabeled datasets for semi-supervised text classification; (2) TestNUC, which improves the test-time performance of large language models (LLMs) by leveraging neighboring unlabeled data consistency without requiring additional labels; and (3) GLEAN, which combines small language models with high-quality feedback generated by LLMs to efficiently generalize category discovery tasks. Building on these insights, we further present a system and benchmark designed to evaluate how effectively and efficiently LLMs can incorporate limited user interruption updates during long-horizon agentic tasks. Collectively, these works chart a path toward building efficient, effective, and scalable AI systems by harmonizing limited labels, model capabilities, and the vast availability of unlabeled data. Experimental evaluations are conducted on publicly available datasets, including AG News, Yahoo! Answers, IMDB, Empathetic Dialogues, GoEmotions, Hurricane, BANKING, CLINC, Reddit, StackExchange, MTOP, FewEvent, StackOverflow and WebArena.","abstract_html":"Efficiently exploiting limited labels alongside abundant unlabeled data is critical for advancing the capabilities of language models and their downstream systems. This thesis first presents novel methodologies for label-efficient learning, introducing three frameworks: (1) JointMatch, which employs collaborative pseudo-labeling to leverage limited labeled examples together with extensive unlabeled datasets for semi-supervised text classification; (2) TestNUC, which improves the test-time performance of large language models (LLMs) by leveraging neighboring unlabeled data consistency without requiring additional labels; and (3) GLEAN, which combines small language models with high-quality feedback generated by LLMs to efficiently generalize category discovery tasks. Building on these insights, we further present a system and benchmark designed to evaluate how effectively and efficiently LLMs can incorporate limited user interruption updates during long-horizon agentic tasks. Collectively, these works chart a path toward building efficient, effective, and scalable AI systems by harmonizing limited labels, model capabilities, and the vast availability of unlabeled data. Experimental evaluations are conducted on publicly available datasets, including AG News, Yahoo! Answers, IMDB, Empathetic Dialogues, GoEmotions, Hurricane, BANKING, CLINC, Reddit, StackExchange, MTOP, FewEvent, StackOverflow and WebArena.","abstract_has_math":false,"creators":["Henry Peng Zou (24399515)"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2026,"date_issued":"2026-05-01T00:00:00Z","date_published":"2026-05-01T00:00:00Z","updated_at":"2026-07-27T21:33:41Z","subjects":["Computer Science","Machine Learning","Artificial Intelligence"],"languages":[],"rights":["In Copyright"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.25417/uic.32993957.v1","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Henry Peng Zou (24399515)"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2026-05-01T00:00:00Z"]},{"key":"dc:relation","label":"Dc Relation","values":["https://figshare.com/articles/thesis/Towards_Label-Efficient_Learning_Exploiting_Limited_Labels_and_Abundant_Unlabeled_Data/32993957"]},{"key":"dc:type","label":"Dc Type","values":["Text","Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science","Machine Learning","Artificial Intelligence"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["10.25417/uic.32993957.v1"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Efficiently exploiting limited labels alongside abundant unlabeled data is critical for advancing the capabilities of language models and their downstream systems. This thesis first presents novel methodologies for label-efficient learning, introducing three frameworks: (1) JointMatch, which employs collaborative pseudo-labeling to leverage limited labeled examples together with extensive unlabeled datasets for semi-supervised text classification; (2) TestNUC, which improves the test-time performance of large language models (LLMs) by leveraging neighboring unlabeled data consistency without requiring additional labels; and (3) GLEAN, which combines small language models with high-quality feedback generated by LLMs to efficiently generalize category discovery tasks. Building on these insights, we further present a system and benchmark designed to evaluate how effectively and efficiently LLMs can incorporate limited user interruption updates during long-horizon agentic tasks. Collectively, these works chart a path toward building efficient, effective, and scalable AI systems by harmonizing limited labels, model capabilities, and the vast availability of unlabeled data. Experimental evaluations are conducted on publicly available datasets, including AG News, Yahoo! Answers, IMDB, Empathetic Dialogues, GoEmotions, Hurricane, BANKING, CLINC, Reddit, StackExchange, MTOP, FewEvent, StackOverflow and WebArena."]},{"key":"dc:title","label":"Title","values":["Towards Label-Efficient Learning: Exploiting Limited Labels and Abundant Unlabeled Data"]}]}],"canonical_facts":{"dc:creator":["Henry Peng Zou (24399515)"],"dc:date":["2026-05-01T00:00:00Z"],"dc:description":["Efficiently exploiting limited labels alongside abundant unlabeled data is critical for advancing the capabilities of language models and their downstream systems. This thesis first presents novel methodologies for label-efficient learning, introducing three frameworks: (1) JointMatch, which employs collaborative pseudo-labeling to leverage limited labeled examples together with extensive unlabeled datasets for semi-supervised text classification; (2) TestNUC, which improves the test-time performance of large language models (LLMs) by leveraging neighboring unlabeled data consistency without requiring additional labels; and (3) GLEAN, which combines small language models with high-quality feedback generated by LLMs to efficiently generalize category discovery tasks. Building on these insights, we further present a system and benchmark designed to evaluate how effectively and efficiently LLMs can incorporate limited user interruption updates during long-horizon agentic tasks. Collectively, these works chart a path toward building efficient, effective, and scalable AI systems by harmonizing limited labels, model capabilities, and the vast availability of unlabeled data. Experimental evaluations are conducted on publicly available datasets, including AG News, Yahoo! Answers, IMDB, Empathetic Dialogues, GoEmotions, Hurricane, BANKING, CLINC, Reddit, StackExchange, MTOP, FewEvent, StackOverflow and WebArena."],"dc:identifier":["10.25417/uic.32993957.v1"],"dc:relation":["https://figshare.com/articles/thesis/Towards_Label-Efficient_Learning_Exploiting_Limited_Labels_and_Abundant_Unlabeled_Data/32993957"],"dc:rights":["In Copyright"],"dc:subject":["Computer Science","Machine Learning","Artificial Intelligence"],"dc:title":["Towards Label-Efficient Learning: Exploiting Limited Labels and Abundant Unlabeled Data"],"dc:type":["Text","Thesis"]},"updated_at":"2026-07-27T21:33:41Z"}