{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/92868"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/92868","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Towards open-ended crowd-powered data processing: a case study of clustering and counting","abstract":"Due to the widespread use and importance of crowdsourcing in gathering training data at scale, the data management community has devoted its efforts in understanding and optimizing fundamental primitives like filters and joins. These primitive boolean operations, where the human responses come from a small, finite space of possible answers, are inadequate for a number of data analysis tasks, especially those involving images, videos and maps. There is, thus, a need for open-ended crowdsourcing in order to get more fine-grained information from humans that can be used in developing sophisticated AI systems. In this thesis, we study two popular open-ended crowdsourcing problems. The first, clustering, is the problem of organizing a collection of objects (images, videos) by allowing workers to form as many clusters as they would like and organize items across them. The second, counting, is the problem of counting objects in images. In this thesis, we develop models to reason about human behavior for both problems, and use these models to design provably cost-efficient algorithms that provide high-quality results, as compared to currently available approaches.","abstract_html":"Due to the widespread use and importance of crowdsourcing in gathering training data at scale, the data management community has devoted its efforts in understanding and optimizing fundamental primitives like filters and joins. These primitive boolean operations, where the human responses come from a small, finite space of possible answers, are inadequate for a number of data analysis tasks, especially those involving images, videos and maps. There is, thus, a need for open-ended crowdsourcing in order to get more fine-grained information from humans that can be used in developing sophisticated AI systems. In this thesis, we study two popular open-ended crowdsourcing problems. The first, clustering, is the problem of organizing a collection of objects (images, videos) by allowing workers to form as many clusters as they would like and organize items across them. The second, counting, is the problem of counting objects in images. In this thesis, we develop models to reason about human behavior for both problems, and use these models to design provably cost-efficient algorithms that provide high-quality results, as compared to currently available approaches.","abstract_has_math":false,"creators":["Jain, Ayush"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Parameswaran, Aditya G."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-11-10T17:55:19Z","date_published":"2016-11-10T17:55:19Z","updated_at":"2026-07-22T22:26:35Z","subjects":["crowdsourcing"],"languages":["en"],"rights":["Copyright 2016 Ayush Jain"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/92868","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Parameswaran, Aditya G."]},{"key":"dc:creator","label":"Author","values":["Jain, Ayush"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-11-10T17:55:19Z","2016-07-19","2016-08"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["crowdsourcing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Ayush Jain"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/92868"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Due to the widespread use and importance of crowdsourcing in gathering training data at scale, the data management community has devoted its efforts in understanding and optimizing fundamental primitives like filters and joins. These primitive boolean operations, where the human responses come from a small, finite space of possible answers, are inadequate for a number of data analysis tasks, especially those involving images, videos and maps. There is, thus, a need for open-ended crowdsourcing in order to get more fine-grained information from humans that can be used in developing sophisticated AI systems. In this thesis, we study two popular open-ended crowdsourcing problems. The first, clustering, is the problem of organizing a collection of objects (images, videos) by allowing workers to form as many clusters as they would like and organize items across them. The second, counting, is the problem of counting objects in images. In this thesis, we develop models to reason about human behavior for both problems, and use these models to design provably cost-efficient algorithms that provide high-quality results, as compared to currently available approaches.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Ayush Jain, accepted the attached license on 2016-07-18 at 16:26.","The student, Ayush Jain, submitted this Thesis for approval on 2016-07-18 at 16:39.","This Thesis was approved for publication on 2016-07-19 at 09:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9998 on 2016-11-09 at 10:25:26","Made available in DSpace on 2016-11-10T17:55:19Z (GMT). No. of bitstreams: 2 JAIN-THESIS-2016.pdf: 18941114 bytes, checksum: 510809166171ffa8f59f0e8aee01d958 (MD5) LICENSE.txt: 4207 bytes, checksum: ab36553fc08758b3f1fc1cb0e7ffcf5b (MD5) Previous issue date: 2016-07-19"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Towards open-ended crowd-powered data processing: a case study of clustering and counting"]}]}],"canonical_facts":{"dc:contributor":["Parameswaran, Aditya G."],"dc:creator":["Jain, Ayush"],"dc:date":["2016-11-10T17:55:19Z","2016-07-19","2016-08"],"dc:description":["Due to the widespread use and importance of crowdsourcing in gathering training data at scale, the data management community has devoted its efforts in understanding and optimizing fundamental primitives like filters and joins. These primitive boolean operations, where the human responses come from a small, finite space of possible answers, are inadequate for a number of data analysis tasks, especially those involving images, videos and maps. There is, thus, a need for open-ended crowdsourcing in order to get more fine-grained information from humans that can be used in developing sophisticated AI systems. In this thesis, we study two popular open-ended crowdsourcing problems. The first, clustering, is the problem of organizing a collection of objects (images, videos) by allowing workers to form as many clusters as they would like and organize items across them. The second, counting, is the problem of counting objects in images. In this thesis, we develop models to reason about human behavior for both problems, and use these models to design provably cost-efficient algorithms that provide high-quality results, as compared to currently available approaches.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-11-09 without embargo terms","The student, Ayush Jain, accepted the attached license on 2016-07-18 at 16:26.","The student, Ayush Jain, submitted this Thesis for approval on 2016-07-18 at 16:39.","This Thesis was approved for publication on 2016-07-19 at 09:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #9998 on 2016-11-09 at 10:25:26","Made available in DSpace on 2016-11-10T17:55:19Z (GMT). No. of bitstreams: 2 JAIN-THESIS-2016.pdf: 18941114 bytes, checksum: 510809166171ffa8f59f0e8aee01d958 (MD5) LICENSE.txt: 4207 bytes, checksum: ab36553fc08758b3f1fc1cb0e7ffcf5b (MD5) Previous issue date: 2016-07-19"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/92868"],"dc:language":["en"],"dc:rights":["Copyright 2016 Ayush Jain"],"dc:subject":["crowdsourcing"],"dc:title":["Towards open-ended crowd-powered data processing: a case study of clustering and counting"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:35Z"}