{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/46856"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/46856","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Semi-supervised learning and relevance search on networked data","abstract":"Real-world data entities are often connected by meaningful relationships, forming large-scale networks. With the rapid growth of social networks and online relational data, it is widely recognized that networked data are playing increasingly important roles in people's daily life. Based on whether the nodes and edges have different semantic meanings or not, networks can be roughly categorized into heterogeneous and homogeneous networks. Although homogeneous networks have been studied for decades, some problems still remain unsolved. Heterogeneous networks are much more complicated than homogeneous networks, and have not been explored until recently. Therefore, effective and principled algorithms for mining both homogeneous and heterogeneous networks are in great demand. In this thesis, two important and closely related problems, semi-supervised learning and relevance search, are studied on both homogeneous and heterogeneous networks. Different from many existing models, algorithms developed in this thesis are theoretically reasonable, widely applicable with minimum constraints, and provide more informative mining results. First, a label selection criterion is proposed to improve the effectiveness of existing semi-supervised learning models on networks. Second, ranking and semi-supervised learning are integrated together to improve the informativeness of the results. Third, a relevance search algorithm that fully considers the geometric structure of the homogeneous networked data is designed. Finally, the relevance search problem between different types of nodes on heterogeneous networks is studied, and the proposed solution is applied on a network constructed from unstructured text data. Research results introduced in this thesis provide advanced principles and the first few steps towards a complete and systematic solution of mining networked data.","abstract_html":"Real-world data entities are often connected by meaningful relationships, forming large-scale networks. With the rapid growth of social networks and online relational data, it is widely recognized that networked data are playing increasingly important roles in people&#x27;s daily life. Based on whether the nodes and edges have different semantic meanings or not, networks can be roughly categorized into heterogeneous and homogeneous networks. Although homogeneous networks have been studied for decades, some problems still remain unsolved. Heterogeneous networks are much more complicated than homogeneous networks, and have not been explored until recently. Therefore, effective and principled algorithms for mining both homogeneous and heterogeneous networks are in great demand. In this thesis, two important and closely related problems, semi-supervised learning and relevance search, are studied on both homogeneous and heterogeneous networks. Different from many existing models, algorithms developed in this thesis are theoretically reasonable, widely applicable with minimum constraints, and provide more informative mining results. First, a label selection criterion is proposed to improve the effectiveness of existing semi-supervised learning models on networks. Second, ranking and semi-supervised learning are integrated together to improve the informativeness of the results. Third, a relevance search algorithm that fully considers the geometric structure of the homogeneous networked data is designed. Finally, the relevance search problem between different types of nodes on heterogeneous networks is studied, and the proposed solution is applied on a network constructed from unstructured text data. Research results introduced in this thesis provide advanced principles and the first few steps towards a complete and systematic solution of mining networked data.","abstract_has_math":false,"creators":["Ji, Ming"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Han, Jiawei","Roth, Dan","Huang, Thomas S.","Chen, Yuguo","Ye, Jieping"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-01-16T18:18:49Z","date_published":"2014-01-16T18:18:49Z","updated_at":"2026-07-22T22:25:38Z","subjects":["Data Mining","Machine Learning","Semi-supervised Learning","Search","Heterogeneous Networks","Graphs"],"languages":["en"],"rights":["Copyright 2013 Ming Ji"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/46856","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Han, Jiawei","Roth, Dan","Huang, Thomas S.","Chen, Yuguo","Ye, Jieping"]},{"key":"dc:creator","label":"Author","values":["Ji, Ming"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-01-16T18:18:49Z","2016-01-16T11:01:39Z","2013-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Data Mining","Machine Learning","Semi-supervised Learning","Search","Heterogeneous Networks","Graphs"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2013 Ming Ji"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/46856"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Real-world data entities are often connected by meaningful relationships, forming large-scale networks. With the rapid growth of social networks and online relational data, it is widely recognized that networked data are playing increasingly important roles in people's daily life. Based on whether the nodes and edges have different semantic meanings or not, networks can be roughly categorized into heterogeneous and homogeneous networks. Although homogeneous networks have been studied for decades, some problems still remain unsolved. Heterogeneous networks are much more complicated than homogeneous networks, and have not been explored until recently. Therefore, effective and principled algorithms for mining both homogeneous and heterogeneous networks are in great demand. In this thesis, two important and closely related problems, semi-supervised learning and relevance search, are studied on both homogeneous and heterogeneous networks. Different from many existing models, algorithms developed in this thesis are theoretically reasonable, widely applicable with minimum constraints, and provide more informative mining results. First, a label selection criterion is proposed to improve the effectiveness of existing semi-supervised learning models on networks. Second, ranking and semi-supervised learning are integrated together to improve the informativeness of the results. Third, a relevance search algorithm that fully considers the geometric structure of the homogeneous networked data is designed. Finally, the relevance search problem between different types of nodes on heterogeneous networks is studied, and the proposed solution is applied on a network constructed from unstructured text data. Research results introduced in this thesis provide advanced principles and the first few steps towards a complete and systematic solution of mining networked data.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-12-05T14:31:22Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Ji_Ming.pdf: 4868037 bytes, checksum: 94352eaed67d698051fd228ff8c28369 (MD5)","Made available in DSpace on 2014-01-16T18:18:49Z (GMT). No. of bitstreams: 2 Ming_Ji.pdf: 4868037 bytes, checksum: 94352eaed67d698051fd228ff8c28369 (MD5) license.txt: 4056 bytes, checksum: 4cc457ae76c1c72678cfeb2a3ced9b6a (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Seth Robbins (robbins.sd@gmail.com) on 2014-01-16T18:19:50Z Item is restricted until 2016-01-16T18:19:34Z","Restriction data tranferred 2014-07-01T11:33:30-05:00 Original Data Group with Access UIUC Users [automated] Release Date: 2016-01-16 12:19:34 UTC Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 46875 on 2016-01-16T11:01:39Z."]},{"key":"dc:title","label":"Title","values":["Semi-supervised learning and relevance search on networked data"]}]}],"canonical_facts":{"dc:contributor":["Han, Jiawei","Roth, Dan","Huang, Thomas S.","Chen, Yuguo","Ye, Jieping"],"dc:creator":["Ji, Ming"],"dc:date":["2014-01-16T18:18:49Z","2016-01-16T11:01:39Z","2013-12"],"dc:description":["Real-world data entities are often connected by meaningful relationships, forming large-scale networks. With the rapid growth of social networks and online relational data, it is widely recognized that networked data are playing increasingly important roles in people's daily life. Based on whether the nodes and edges have different semantic meanings or not, networks can be roughly categorized into heterogeneous and homogeneous networks. Although homogeneous networks have been studied for decades, some problems still remain unsolved. Heterogeneous networks are much more complicated than homogeneous networks, and have not been explored until recently. Therefore, effective and principled algorithms for mining both homogeneous and heterogeneous networks are in great demand. In this thesis, two important and closely related problems, semi-supervised learning and relevance search, are studied on both homogeneous and heterogeneous networks. Different from many existing models, algorithms developed in this thesis are theoretically reasonable, widely applicable with minimum constraints, and provide more informative mining results. First, a label selection criterion is proposed to improve the effectiveness of existing semi-supervised learning models on networks. Second, ranking and semi-supervised learning are integrated together to improve the informativeness of the results. Third, a relevance search algorithm that fully considers the geometric structure of the homogeneous networked data is designed. Finally, the relevance search problem between different types of nodes on heterogeneous networks is studied, and the proposed solution is applied on a network constructed from unstructured text data. Research results introduced in this thesis provide advanced principles and the first few steps towards a complete and systematic solution of mining networked data.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2013-12-05T14:31:22Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Ji_Ming.pdf: 4868037 bytes, checksum: 94352eaed67d698051fd228ff8c28369 (MD5)","Made available in DSpace on 2014-01-16T18:18:49Z (GMT). No. of bitstreams: 2 Ming_Ji.pdf: 4868037 bytes, checksum: 94352eaed67d698051fd228ff8c28369 (MD5) license.txt: 4056 bytes, checksum: 4cc457ae76c1c72678cfeb2a3ced9b6a (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by Seth Robbins (robbins.sd@gmail.com) on 2014-01-16T18:19:50Z Item is restricted until 2016-01-16T18:19:34Z","Restriction data tranferred 2014-07-01T11:33:30-05:00 Original Data Group with Access UIUC Users [automated] Release Date: 2016-01-16 12:19:34 UTC Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 46875 on 2016-01-16T11:01:39Z."],"dc:identifier":["http://hdl.handle.net/2142/46856"],"dc:language":["en"],"dc:rights":["Copyright 2013 Ming Ji"],"dc:subject":["Data Mining","Machine Learning","Semi-supervised Learning","Search","Heterogeneous Networks","Graphs"],"dc:title":["Semi-supervised learning and relevance search on networked data"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:38Z"}