{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/115353"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/115353","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Optimizing the wisdom of the crowd: Learning, teaching, and recommendation","abstract":"Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_html":"Submission original under an indefinite embargo labeled &#x27;Open Access&#x27;. The submission was exported from vireo on 2022-11-11 without embargo terms","abstract_has_math":false,"creators":["Zhou, Yao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["He, Jingrui","Zhai, Chengxiang","Li, Bo","Ying, Lei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-05","date_published":"2022-05","updated_at":"2026-07-22T22:24:54Z","subjects":["Crowdsourcing","Wisdom of crowd","Truth Inference","Heterogeneous Learning","Machine Teaching","Recommender Systems"],"languages":["en","eng"],"rights":["Copyright 2022 Yao Zhou"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/115353","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["He, Jingrui","Zhai, Chengxiang","Li, Bo","Ying, Lei"]},{"key":"dc:creator","label":"Author","values":["Zhou, Yao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-05","2022-04-06"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Crowdsourcing","Wisdom of crowd","Truth Inference","Heterogeneous Learning","Machine Teaching","Recommender Systems"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Yao Zhou"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/115353"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Yao Zhou, accepted the attached license on 2022-04-06 at 02:14.","The student, Yao Zhou, submitted this Dissertation for approval on 2022-04-06 at 02:26.","This Dissertation was approved for publication on 2022-04-06 at 16:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17545 on 2022-11-11 at 13:04:35","Crowdsourcing is one type of sourcing model that helps individuals and organizations obtain large amount of information, e.g., ideas, knowledge, financial resources, etc., from a group of non-experts within limited time at a low cost. In recent decades, the unprecedented amount of data has catalyzed the trend of combining human insights with machine learning techniques, and thereby, facilitated the wide usage of crowdsourcing to enlist label information effectively and efficiently. Specifically, the adoption of crowdsourcing has been witnessed in a wide range of high-impact real-world applications, such as collaborative knowledge (e.g., data annotation, language translation), collective creativity (e.g., analogy mining, crowdfunding), and reverse Turing test (e.g., CAPTCHA-like systems), etc. Researchers and practitioners have mainly worked on the collaborative information aggregation problem where the data items are outsourced to a group of unskilled online workers for labeling. Despite the wide adoption of crowdsourcing services, several of its fundamental problems remain unsolved especially at the information and cognitive levels in terms of incentive design, aggregation methodology, and model learning. Specifically, the related computational challenges fall into the following categories of problems: (P1) Truth inference - How to infer the ground truth label of each item that has been distributed and labeled by multiple crowdsourcing labelers? What is the benefit of using truth inference for the downstream learning tasks? (P2) Heterogeneous learning - How to perform a unified model learning when the labeled data demonstrate multiple types of heterogeneity, such as oracle heterogeneity (more than one labels are collected from different labelers per item), task heterogeneity (multiple similar predictive tasks share commonality and are jointly learned), and view heterogeneity (items are characterized by different sources of features)? (P3) Crowd teaching - How to improve the labeling expertise of the crowdsourcing workers in the form of teaching and as a response, benefit the model learning eventually? The teaching mechanism should be personalized, effective, and interpretable. (P4) Recommendation - How to rank the user relevance over the unseen items if only a small number of user’s ratings (instead of the class labels in crowdsourcing) over items are given? How to properly handle the data sparsity, imbalanced sampling, and provide explanations to justify the recommendations? In this thesis, an in-depth study of these four problems is properly motivated, rigorously formulated, thoroughly evaluated, and broadly discussed. Specifically, one tensor augmentation and completion model is proposed to tackle the problem of Truth Inference; two unified optimization frameworks are proposed to solve the Heterogeneous Learning problem with crowdsourced labels that have various types of data heterogeneities - task heterogeneity and view heterogeneity; two Crowd Teaching mechanisms are proposed to guide crowd workers toward the personalized teaching as well as the interpretable concept learning. In the end, we tackle the Recommendation problem by proposing one unified ranking model that combines positive-unlabeled learning with generative adversarial networks to mitigate the data sparsity in collaborative filtering and also one systematical study regarding the explainability of the contextualized recommender systems. Optimizing the wisdom of the crowd is a promising direction with lots of potentials. Towards the end of the thesis, we have also presented a few promising future research directions such as complementary recommendations, fairness recommendations, and federated learning."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Optimizing the wisdom of the crowd: Learning, teaching, and recommendation"]}]}],"canonical_facts":{"dc:contributor":["He, Jingrui","Zhai, Chengxiang","Li, Bo","Ying, Lei"],"dc:creator":["Zhou, Yao"],"dc:date":["2022-05","2022-04-06"],"dc:description":["Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-11-11 without embargo terms","The student, Yao Zhou, accepted the attached license on 2022-04-06 at 02:14.","The student, Yao Zhou, submitted this Dissertation for approval on 2022-04-06 at 02:26.","This Dissertation was approved for publication on 2022-04-06 at 16:01.","DSpace SAF Submission Ingestion Package generated from Vireo submission #17545 on 2022-11-11 at 13:04:35","Crowdsourcing is one type of sourcing model that helps individuals and organizations obtain large amount of information, e.g., ideas, knowledge, financial resources, etc., from a group of non-experts within limited time at a low cost. In recent decades, the unprecedented amount of data has catalyzed the trend of combining human insights with machine learning techniques, and thereby, facilitated the wide usage of crowdsourcing to enlist label information effectively and efficiently. Specifically, the adoption of crowdsourcing has been witnessed in a wide range of high-impact real-world applications, such as collaborative knowledge (e.g., data annotation, language translation), collective creativity (e.g., analogy mining, crowdfunding), and reverse Turing test (e.g., CAPTCHA-like systems), etc. Researchers and practitioners have mainly worked on the collaborative information aggregation problem where the data items are outsourced to a group of unskilled online workers for labeling. Despite the wide adoption of crowdsourcing services, several of its fundamental problems remain unsolved especially at the information and cognitive levels in terms of incentive design, aggregation methodology, and model learning. Specifically, the related computational challenges fall into the following categories of problems: (P1) Truth inference - How to infer the ground truth label of each item that has been distributed and labeled by multiple crowdsourcing labelers? What is the benefit of using truth inference for the downstream learning tasks? (P2) Heterogeneous learning - How to perform a unified model learning when the labeled data demonstrate multiple types of heterogeneity, such as oracle heterogeneity (more than one labels are collected from different labelers per item), task heterogeneity (multiple similar predictive tasks share commonality and are jointly learned), and view heterogeneity (items are characterized by different sources of features)? (P3) Crowd teaching - How to improve the labeling expertise of the crowdsourcing workers in the form of teaching and as a response, benefit the model learning eventually? The teaching mechanism should be personalized, effective, and interpretable. (P4) Recommendation - How to rank the user relevance over the unseen items if only a small number of user’s ratings (instead of the class labels in crowdsourcing) over items are given? How to properly handle the data sparsity, imbalanced sampling, and provide explanations to justify the recommendations? In this thesis, an in-depth study of these four problems is properly motivated, rigorously formulated, thoroughly evaluated, and broadly discussed. Specifically, one tensor augmentation and completion model is proposed to tackle the problem of Truth Inference; two unified optimization frameworks are proposed to solve the Heterogeneous Learning problem with crowdsourced labels that have various types of data heterogeneities - task heterogeneity and view heterogeneity; two Crowd Teaching mechanisms are proposed to guide crowd workers toward the personalized teaching as well as the interpretable concept learning. In the end, we tackle the Recommendation problem by proposing one unified ranking model that combines positive-unlabeled learning with generative adversarial networks to mitigate the data sparsity in collaborative filtering and also one systematical study regarding the explainability of the contextualized recommender systems. Optimizing the wisdom of the crowd is a promising direction with lots of potentials. Towards the end of the thesis, we have also presented a few promising future research directions such as complementary recommendations, fairness recommendations, and federated learning."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/115353"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Yao Zhou"],"dc:subject":["Crowdsourcing","Wisdom of crowd","Truth Inference","Heterogeneous Learning","Machine Teaching","Recommender Systems"],"dc:title":["Optimizing the wisdom of the crowd: Learning, teaching, and recommendation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:54Z"}