{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/95303"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/95303","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Similarity learning in the era of big data","abstract":"This dissertation studies the problem of similarity learning in the era of big data with heavy emphasis on real-world applications in social media. As in the saying “birds of a feather flock together,” in similarity learning, we aim to identify the notion of being similar in a data-driven and task-specific way, which is a central problem for maximizing the value of big data. Despite many successes of similarity learning from past decades, social media networks as one of the most typical big data media contain large-volume, various and high-velocity data, which makes conventional learning paradigms and off- the-shelf algorithms insufficient. Thus, we focus on addressing the emerging challenges brought by the inherent “three-Vs” characteristics of big data by answering the following questions: 1) Similarity is characterized by both links and node contents in networks; how to identify the contribution of each network component to seamlessly construct an application orientated similarity function? 2) Social media data are massive and contain much noise; how to efficiently learn the similarity between node pairs in large and noisy environments? 3) Node contents in social media networks are multi-modal; how to effectively measure cross-modal similarity by bridging the so-called “semantic gap”? 4) User wants and needs, and item characteristics, are continuously evolving, which generates data at an unprecedented rate; how to model the nature of temporal dynamics in principle and provide timely decision makings? The goal of this dissertation is to provide solutions to these questions via innovative research and novel methods. We hope this dissertation sheds more light on similarity learning in the big data era and broadens its applications in social media.","abstract_html":"This dissertation studies the problem of similarity learning in the era of big data with heavy emphasis on real-world applications in social media. As in the saying “birds of a feather flock together,” in similarity learning, we aim to identify the notion of being similar in a data-driven and task-specific way, which is a central problem for maximizing the value of big data. Despite many successes of similarity learning from past decades, social media networks as one of the most typical big data media contain large-volume, various and high-velocity data, which makes conventional learning paradigms and off- the-shelf algorithms insufficient. Thus, we focus on addressing the emerging challenges brought by the inherent “three-Vs” characteristics of big data by answering the following questions: 1) Similarity is characterized by both links and node contents in networks; how to identify the contribution of each network component to seamlessly construct an application orientated similarity function? 2) Social media data are massive and contain much noise; how to efficiently learn the similarity between node pairs in large and noisy environments? 3) Node contents in social media networks are multi-modal; how to effectively measure cross-modal similarity by bridging the so-called “semantic gap”? 4) User wants and needs, and item characteristics, are continuously evolving, which generates data at an unprecedented rate; how to model the nature of temporal dynamics in principle and provide timely decision makings? The goal of this dissertation is to provide solutions to these questions via innovative research and novel methods. We hope this dissertation sheds more light on similarity learning in the big data era and broadens its applications in social media.","abstract_has_math":false,"creators":["Chang, Shiyu"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Huang, Thomas S.","Hasegawa-Johnson, Mark A.","Liang, Zhi-Pei","Aggarwal, Charu C."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-03-01T15:46:18Z","date_published":"2017-03-01T15:46:18Z","updated_at":"2026-07-22T22:26:37Z","subjects":["Similarity Learning","Big Data","Large Volume Data","Multimodality Data","High-velocity Data","Large-scale","Supervised Similarity Learning","Network Embedding","Deep Embedding","Heterogeneous Network","Streaming Network","Positive-Unlabeled Learning","Link Prediction","Recommendation","Social Media","Search and Retrieval"],"languages":["en"],"rights":["Copyright 2016 Shiyu Chang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/95303","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Thomas S.","Hasegawa-Johnson, Mark A.","Liang, Zhi-Pei","Aggarwal, Charu C."]},{"key":"dc:creator","label":"Author","values":["Chang, Shiyu"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-03-01T15:46:18Z","2016-11-04","2016-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Similarity Learning","Big Data","Large Volume Data","Multimodality Data","High-velocity Data","Large-scale","Supervised Similarity Learning","Network Embedding","Deep Embedding","Heterogeneous Network","Streaming Network","Positive-Unlabeled Learning","Link Prediction","Recommendation","Social Media","Search and Retrieval"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Shiyu Chang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/95303"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation studies the problem of similarity learning in the era of big data with heavy emphasis on real-world applications in social media. As in the saying “birds of a feather flock together,” in similarity learning, we aim to identify the notion of being similar in a data-driven and task-specific way, which is a central problem for maximizing the value of big data. Despite many successes of similarity learning from past decades, social media networks as one of the most typical big data media contain large-volume, various and high-velocity data, which makes conventional learning paradigms and off- the-shelf algorithms insufficient. Thus, we focus on addressing the emerging challenges brought by the inherent “three-Vs” characteristics of big data by answering the following questions: 1) Similarity is characterized by both links and node contents in networks; how to identify the contribution of each network component to seamlessly construct an application orientated similarity function? 2) Social media data are massive and contain much noise; how to efficiently learn the similarity between node pairs in large and noisy environments? 3) Node contents in social media networks are multi-modal; how to effectively measure cross-modal similarity by bridging the so-called “semantic gap”? 4) User wants and needs, and item characteristics, are continuously evolving, which generates data at an unprecedented rate; how to model the nature of temporal dynamics in principle and provide timely decision makings? The goal of this dissertation is to provide solutions to these questions via innovative research and novel methods. We hope this dissertation sheds more light on similarity learning in the big data era and broadens its applications in social media.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-02-28 without embargo terms","The student, Shiyu Chang, accepted the attached license on 2016-10-31 at 14:05.","The student, Shiyu Chang, submitted this Dissertation for approval on 2016-10-31 at 14:27.","This Dissertation was approved for publication on 2016-11-04 at 14:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10211 on 2017-02-28 at 14:46:43","Made available in DSpace on 2017-03-01T15:46:18Z (GMT). No. of bitstreams: 5 CHANG-DISSERTATION-2016.pdf: 3692195 bytes, checksum: ec9f830fcc07055f29734cb20ced8e42 (MD5) LICENSE.txt: 4208 bytes, checksum: 4674edb664b84f701928e466bcc2e0d8 (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: f50bdef2cac3da2764bf77cc65a040b1 (MD5) RightsLink Printable License.pdf: 161443 bytes, checksum: 09792450e193d6c255a379c89ee4dd08 (MD5) RightsLink Printable License_1.pdf: 161443 bytes, checksum: 09792450e193d6c255a379c89ee4dd08 (MD5) Previous issue date: 2016-11-04"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Similarity learning in the era of big data"]}]}],"canonical_facts":{"dc:contributor":["Huang, Thomas S.","Hasegawa-Johnson, Mark A.","Liang, Zhi-Pei","Aggarwal, Charu C."],"dc:creator":["Chang, Shiyu"],"dc:date":["2017-03-01T15:46:18Z","2016-11-04","2016-12"],"dc:description":["This dissertation studies the problem of similarity learning in the era of big data with heavy emphasis on real-world applications in social media. As in the saying “birds of a feather flock together,” in similarity learning, we aim to identify the notion of being similar in a data-driven and task-specific way, which is a central problem for maximizing the value of big data. Despite many successes of similarity learning from past decades, social media networks as one of the most typical big data media contain large-volume, various and high-velocity data, which makes conventional learning paradigms and off- the-shelf algorithms insufficient. Thus, we focus on addressing the emerging challenges brought by the inherent “three-Vs” characteristics of big data by answering the following questions: 1) Similarity is characterized by both links and node contents in networks; how to identify the contribution of each network component to seamlessly construct an application orientated similarity function? 2) Social media data are massive and contain much noise; how to efficiently learn the similarity between node pairs in large and noisy environments? 3) Node contents in social media networks are multi-modal; how to effectively measure cross-modal similarity by bridging the so-called “semantic gap”? 4) User wants and needs, and item characteristics, are continuously evolving, which generates data at an unprecedented rate; how to model the nature of temporal dynamics in principle and provide timely decision makings? The goal of this dissertation is to provide solutions to these questions via innovative research and novel methods. We hope this dissertation sheds more light on similarity learning in the big data era and broadens its applications in social media.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-02-28 without embargo terms","The student, Shiyu Chang, accepted the attached license on 2016-10-31 at 14:05.","The student, Shiyu Chang, submitted this Dissertation for approval on 2016-10-31 at 14:27.","This Dissertation was approved for publication on 2016-11-04 at 14:14.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10211 on 2017-02-28 at 14:46:43","Made available in DSpace on 2017-03-01T15:46:18Z (GMT). No. of bitstreams: 5 CHANG-DISSERTATION-2016.pdf: 3692195 bytes, checksum: ec9f830fcc07055f29734cb20ced8e42 (MD5) LICENSE.txt: 4208 bytes, checksum: 4674edb664b84f701928e466bcc2e0d8 (MD5) PROQUEST_LICENSE.txt: 4554 bytes, checksum: f50bdef2cac3da2764bf77cc65a040b1 (MD5) RightsLink Printable License.pdf: 161443 bytes, checksum: 09792450e193d6c255a379c89ee4dd08 (MD5) RightsLink Printable License_1.pdf: 161443 bytes, checksum: 09792450e193d6c255a379c89ee4dd08 (MD5) Previous issue date: 2016-11-04"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/95303"],"dc:language":["en"],"dc:rights":["Copyright 2016 Shiyu Chang"],"dc:subject":["Similarity Learning","Big Data","Large Volume Data","Multimodality Data","High-velocity Data","Large-scale","Supervised Similarity Learning","Network Embedding","Deep Embedding","Heterogeneous Network","Streaming Network","Positive-Unlabeled Learning","Link Prediction","Recommendation","Social Media","Search and Retrieval"],"dc:title":["Similarity learning in the era of big data"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:37Z"}