{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/14623"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/14623","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Using the Internet for object image retrieval and object image classification","abstract":"The Internet has become the largest repository for numerous resources, a big portion of which are images and related multimedia content such as text and videos. This content is valuable for many computer vision tasks. In this thesis, two case studies are conducted to show how to leverage information from the Internet for two important computer vision tasks: object image retrieval and object image classification. Case study 1 is on object image retrieval. With specified object class labels, we aim to retrieve relevant images found on web pages using an analysis of text around the image and of image appearance. For this task, we exploit established online knowledge resources (Wikipedia pages for text; Flickr and Caltech data sets for images). These resources provide rich text and object appearance information. We describe results on two data sets. The first is Berg’s collection of 10 animal categories; on this data set, we significantly outperform previous approaches. In addition, we have collected 5 more categories, and experimental results also show the effectiveness of our approach on this new data set. Case study 2 is on object image classification. We introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text features of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well; it consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small.","abstract_html":"The Internet has become the largest repository for numerous resources, a big portion of which are images and related multimedia content such as text and videos. This content is valuable for many computer vision tasks. In this thesis, two case studies are conducted to show how to leverage information from the Internet for two important computer vision tasks: object image retrieval and object image classification. Case study 1 is on object image retrieval. With specified object class labels, we aim to retrieve relevant images found on web pages using an analysis of text around the image and of image appearance. For this task, we exploit established online knowledge resources (Wikipedia pages for text; Flickr and Caltech data sets for images). These resources provide rich text and object appearance information. We describe results on two data sets. The first is Berg’s collection of 10 animal categories; on this data set, we significantly outperform previous approaches. In addition, we have collected 5 more categories, and experimental results also show the effectiveness of our approach on this new data set. Case study 2 is on object image classification. We introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text features of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well; it consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small.","abstract_has_math":false,"creators":["Wang, Gang"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Forsyth, David A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-01-06T16:20:00Z","date_published":"2010-01-06T16:20:00Z","updated_at":"2026-07-22T22:25:07Z","subjects":["Internet","Object image retrieval","Object image classification"],"languages":["en"],"rights":["Copyright 2009 Gang Wang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/14623","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David A."]},{"key":"dc:creator","label":"Author","values":["Wang, Gang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2010-01-06T16:20:00Z","2009-12"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Internet","Object image retrieval","Object image classification"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2009 Gang Wang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/14623"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The Internet has become the largest repository for numerous resources, a big portion of which are images and related multimedia content such as text and videos. This content is valuable for many computer vision tasks. In this thesis, two case studies are conducted to show how to leverage information from the Internet for two important computer vision tasks: object image retrieval and object image classification. Case study 1 is on object image retrieval. With specified object class labels, we aim to retrieve relevant images found on web pages using an analysis of text around the image and of image appearance. For this task, we exploit established online knowledge resources (Wikipedia pages for text; Flickr and Caltech data sets for images). These resources provide rich text and object appearance information. We describe results on two data sets. The first is Berg’s collection of 10 animal categories; on this data set, we significantly outperform previous approaches. In addition, we have collected 5 more categories, and experimental results also show the effectiveness of our approach on this new data set. Case study 2 is on object image classification. We introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text features of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well; it consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2009-12-11T17:19:29Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Wang_Gang.pdf: 18219489 bytes, checksum: 9f79219964b1b26f17a3ca0127f570db (MD5)","Made available in DSpace on 2010-01-06T16:20:00Z (GMT). No. of bitstreams: 2 license.txt: 4057 bytes, checksum: eca76611549967cb893ae399dddf61e0 (MD5) Wang_Gang.pdf: 18219489 bytes, checksum: 9f79219964b1b26f17a3ca0127f570db (MD5)"]},{"key":"dc:title","label":"Title","values":["Using the Internet for object image retrieval and object image classification"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David A."],"dc:creator":["Wang, Gang"],"dc:date":["2010-01-06T16:20:00Z","2009-12"],"dc:description":["The Internet has become the largest repository for numerous resources, a big portion of which are images and related multimedia content such as text and videos. This content is valuable for many computer vision tasks. In this thesis, two case studies are conducted to show how to leverage information from the Internet for two important computer vision tasks: object image retrieval and object image classification. Case study 1 is on object image retrieval. With specified object class labels, we aim to retrieve relevant images found on web pages using an analysis of text around the image and of image appearance. For this task, we exploit established online knowledge resources (Wikipedia pages for text; Flickr and Caltech data sets for images). These resources provide rich text and object appearance information. We describe results on two data sets. The first is Berg’s collection of 10 animal categories; on this data set, we significantly outperform previous approaches. In addition, we have collected 5 more categories, and experimental results also show the effectiveness of our approach on this new data set. Case study 2 is on object image classification. We introduce a text-based image feature and demonstrate that it consistently improves performance on hard object classification problems. The feature is built using an auxiliary dataset of images annotated with tags, downloaded from the Internet. We do not inspect or correct the tags and expect that they are noisy. We obtain the text features of an unannotated image from the tags of its k-nearest neighbors in this auxiliary collection. A visual classifier presented with an object viewed under novel circumstances (say, a new viewing direction) must rely on its visual examples. Our text feature may not change, because the auxiliary dataset likely contains a similar picture. While the tags associated with images are noisy, they are more stable when appearance changes. We test the performance of this feature using PASCAL VOC 2006 and 2007 datasets. Our feature performs well; it consistently improves the performance of visual object classifiers, and is particularly effective when the training dataset is small.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2009-12-11T17:19:29Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Wang_Gang.pdf: 18219489 bytes, checksum: 9f79219964b1b26f17a3ca0127f570db (MD5)","Made available in DSpace on 2010-01-06T16:20:00Z (GMT). No. of bitstreams: 2 license.txt: 4057 bytes, checksum: eca76611549967cb893ae399dddf61e0 (MD5) Wang_Gang.pdf: 18219489 bytes, checksum: 9f79219964b1b26f17a3ca0127f570db (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/14623"],"dc:language":["en"],"dc:rights":["Copyright 2009 Gang Wang"],"dc:subject":["Internet","Object image retrieval","Object image classification"],"dc:title":["Using the Internet for object image retrieval and object image classification"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}