{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/104942"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/104942","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Classifying GitHub repositories with minimal human efforts","abstract":"GitHub is a great platform for sharing software code, data, and other resources. To improve search and analysis of a vast spectrum of resources on GitHub, it is necessary to conduct automatic, flexible and user-guided classification of GitHub repositories. In this paper, we study how to build a customized repository classifier with minimal human annotation. Previous document classification methods cannot be directly applied to our task due to three unique challenges: (1) multi-modal signals: besides text, signals in other formats need to be explored to uncover the topic of a repository; (2) low data quality: GitHub README files, usually containing code and commands, are noisier than typical text data such as scientific papers and news; and (3) limited ground-truth: users cannot afford to label many repositories for training a good classifier. To deal with the challenges above, we propose GitClass, a framework to classify GitHub repositories under weak supervision. Three key modules, heterogeneous network construction and embedding, keyword extraction and topic modeling, as well as pseudo document generation, are used to tackle the above three challenges, respectively. We conduct extensive experiments on three large-scale GitHub repository datasets and observe evident performance boost over state-of-the-art embedding and classification algorithms.","abstract_html":"GitHub is a great platform for sharing software code, data, and other resources. To improve search and analysis of a vast spectrum of resources on GitHub, it is necessary to conduct automatic, flexible and user-guided classification of GitHub repositories. In this paper, we study how to build a customized repository classifier with minimal human annotation. Previous document classification methods cannot be directly applied to our task due to three unique challenges: (1) multi-modal signals: besides text, signals in other formats need to be explored to uncover the topic of a repository; (2) low data quality: GitHub README files, usually containing code and commands, are noisier than typical text data such as scientific papers and news; and (3) limited ground-truth: users cannot afford to label many repositories for training a good classifier. To deal with the challenges above, we propose GitClass, a framework to classify GitHub repositories under weak supervision. Three key modules, heterogeneous network construction and embedding, keyword extraction and topic modeling, as well as pseudo document generation, are used to tackle the above three challenges, respectively. We conduct extensive experiments on three large-scale GitHub repository datasets and observe evident performance boost over state-of-the-art embedding and classification algorithms.","abstract_has_math":false,"creators":["Zhang, Yu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-08-23T20:05:22Z","date_published":"2019-08-23T20:05:22Z","updated_at":"2026-07-22T22:24:42Z","subjects":["GitHub","Classification","Weak Supervision"],"languages":["en"],"rights":["Copyright 2019 Yu Zhang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/104942","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei"]},{"key":"dc:creator","label":"Author","values":["Zhang, Yu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-08-23T20:05:22Z","2019-04-26","2019-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["GitHub","Classification","Weak Supervision"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Yu Zhang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/104942"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["GitHub is a great platform for sharing software code, data, and other resources. To improve search and analysis of a vast spectrum of resources on GitHub, it is necessary to conduct automatic, flexible and user-guided classification of GitHub repositories. In this paper, we study how to build a customized repository classifier with minimal human annotation. Previous document classification methods cannot be directly applied to our task due to three unique challenges: (1) multi-modal signals: besides text, signals in other formats need to be explored to uncover the topic of a repository; (2) low data quality: GitHub README files, usually containing code and commands, are noisier than typical text data such as scientific papers and news; and (3) limited ground-truth: users cannot afford to label many repositories for training a good classifier. To deal with the challenges above, we propose GitClass, a framework to classify GitHub repositories under weak supervision. Three key modules, heterogeneous network construction and embedding, keyword extraction and topic modeling, as well as pseudo document generation, are used to tackle the above three challenges, respectively. We conduct extensive experiments on three large-scale GitHub repository datasets and observe evident performance boost over state-of-the-art embedding and classification algorithms.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Yu Zhang, accepted the attached license on 2019-04-25 at 16:07.","The student, Yu Zhang, submitted this Thesis for approval on 2019-04-25 at 16:12.","This Thesis was approved for publication on 2019-04-26 at 09:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13926 on 2019-08-22 at 14:46:52","Made available in DSpace on 2019-08-23T20:05:22Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2019.pdf: 4304238 bytes, checksum: 0aea0722ad2fffdedc679674e52f0aad (MD5) LICENSE.txt: 4205 bytes, checksum: d9b2a8b2c7fa1216828f5acea19fdaf4 (MD5) Previous issue date: 2019-04-26"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Classifying GitHub repositories with minimal human efforts"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei"],"dc:creator":["Zhang, Yu"],"dc:date":["2019-08-23T20:05:22Z","2019-04-26","2019-05"],"dc:description":["GitHub is a great platform for sharing software code, data, and other resources. To improve search and analysis of a vast spectrum of resources on GitHub, it is necessary to conduct automatic, flexible and user-guided classification of GitHub repositories. In this paper, we study how to build a customized repository classifier with minimal human annotation. Previous document classification methods cannot be directly applied to our task due to three unique challenges: (1) multi-modal signals: besides text, signals in other formats need to be explored to uncover the topic of a repository; (2) low data quality: GitHub README files, usually containing code and commands, are noisier than typical text data such as scientific papers and news; and (3) limited ground-truth: users cannot afford to label many repositories for training a good classifier. To deal with the challenges above, we propose GitClass, a framework to classify GitHub repositories under weak supervision. Three key modules, heterogeneous network construction and embedding, keyword extraction and topic modeling, as well as pseudo document generation, are used to tackle the above three challenges, respectively. We conduct extensive experiments on three large-scale GitHub repository datasets and observe evident performance boost over state-of-the-art embedding and classification algorithms.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2019-08-22 without embargo terms","The student, Yu Zhang, accepted the attached license on 2019-04-25 at 16:07.","The student, Yu Zhang, submitted this Thesis for approval on 2019-04-25 at 16:12.","This Thesis was approved for publication on 2019-04-26 at 09:44.","DSpace SAF Submission Ingestion Package generated from Vireo submission #13926 on 2019-08-22 at 14:46:52","Made available in DSpace on 2019-08-23T20:05:22Z (GMT). No. of bitstreams: 2 ZHANG-THESIS-2019.pdf: 4304238 bytes, checksum: 0aea0722ad2fffdedc679674e52f0aad (MD5) LICENSE.txt: 4205 bytes, checksum: d9b2a8b2c7fa1216828f5acea19fdaf4 (MD5) Previous issue date: 2019-04-26"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/104942"],"dc:language":["en"],"dc:rights":["Copyright 2019 Yu Zhang"],"dc:subject":["GitHub","Classification","Weak Supervision"],"dc:title":["Classifying GitHub repositories with minimal human efforts"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:42Z"}