University of Illinois - Chicago
Towards Label-Efficient Learning: Exploiting Limited Labels and Abundant Unlabeled Data
Abstract
dc:descriptionEfficiently exploiting limited labels alongside abundant unlabeled data is critical for advancing the capabilities of language models and their downstream systems. This thesis first presents novel methodologies for label-efficient learning, introducing three frameworks: (1) JointMatch, which employs collaborative pseudo-labeling to leverage limited labeled examples together with extensive unlabeled datasets for semi-supervised text classification; (2) TestNUC, which improves the test-time performance of large language models (LLMs) by leveraging neighboring unlabeled data consistency without requiring additional labels; and (3) GLEAN, which combines small language models with high-quality feedback generated by LLMs to efficiently generalize category discovery tasks. Building on these insights, we further present a system and benchmark designed to evaluate how effectively and efficiently LLMs can incorporate limited user interruption updates during long-horizon agentic tasks. Collectively, these works chart a path toward building efficient, effective, and scalable AI systems by harmonizing limited labels, model capabilities, and the vast availability of unlabeled data. Experimental evaluations are conducted on publicly available datasets, including AG News, Yahoo! Answers, IMDB, Empathetic Dialogues, GoEmotions, Hurricane, BANKING, CLINC, Reddit, StackExchange, MTOP, FewEvent, StackOverflow and WebArena.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Henry Peng Zou (24399515)
Subjects
dc:subject × 3Rights
dc:rights- Statement dc:rights
-
- In Copyright
Identifiers
dc:identifier.*- DOI dc:identifier
- https://doi.org/10.25417/uic.32993957.v1
- OAI identifier oai:identifier
- oai:figshare.com:article/32993957