{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/32060"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/32060","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Recognition using visual phrases","abstract":"Made available in DSpace on 2012-06-27T21:31:00Z (GMT). No. of bitstreams: 2 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5) license.txt: 4066 bytes, checksum: e11c5bd6828b3d05964feec70e3062c8 (MD5)","abstract_html":"Made available in DSpace on 2012-06-27T21:31:00Z (GMT). No. of bitstreams: 2 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5) license.txt: 4066 bytes, checksum: e11c5bd6828b3d05964feec70e3062c8 (MD5)","abstract_has_math":false,"creators":["Sadeghi, Mohammad Amin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David A."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-06-27T21:31:00Z","date_published":"2012-06-27T21:31:00Z","updated_at":"2026-07-22T22:25:30Z","subjects":["Visual Phrase","Phrasal Recognition","Visual Composites","Object Recognition"],"languages":["en"],"rights":["Copyright 2012 Mohammad Amin Sadeghi and Ali Farhadi under Creative Commons"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/32060","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David A."]},{"key":"dc:creator","label":"Author","values":["Sadeghi, Mohammad Amin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2012-06-27T21:31:00Z","2014-06-28T10:00:28Z","2012-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Visual Phrase","Phrasal Recognition","Visual Composites","Object Recognition"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2012 Mohammad Amin Sadeghi and Ali Farhadi under Creative Commons"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/32060"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Made available in DSpace on 2012-06-27T21:31:00Z (GMT). No. of bitstreams: 2 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5) license.txt: 4066 bytes, checksum: e11c5bd6828b3d05964feec70e3062c8 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by William Ingram (wingram2@illinois.edu) on 2012-06-27T21:32:46Z Item is restricted until 2014-06-27T21:32:23Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:28Z Item was in collections: Graduate Theses and Dissertations at Illinois (ID: 204) Dissertations and Theses - Computer Science (ID: 587) No. of bitstreams: 2 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5) license.txt: 4066 bytes, checksum: e11c5bd6828b3d05964feec70e3062c8 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:28Z","In this thesis I introduce visual phrases, complex visual composites like ``a person riding a horse''. Visual phrases often display significantly reduced visual complexity compared to their component objects, because the appearance of those objects can change profoundly when they participate in relations. I introduce a dataset suitable for phrasal recognition that uses familiar PASCAL object categories, and demonstrate significant experimental gains resulting from exploiting visual phrases. I show that a visual phrase detector significantly outperforms a baseline which detects component objects and reasons about relations, even though visual phrase training sets tend to be smaller than those for objects. I argue that any multi-class detection system must decode detector outputs to produce final results; this is usually done with non-maximum suppression. I describe a novel decoding procedure that can account accurately for local context without solving difficult inference problems. I show this decoding procedure outperforms the state of the art. Finally, I show that decoding a combination of phrasal and object detectors produces real improvements in detector results.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-04-25T20:23:32Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5)"]},{"key":"dc:title","label":"Title","values":["Recognition using visual phrases"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David A."],"dc:creator":["Sadeghi, Mohammad Amin"],"dc:date":["2012-06-27T21:31:00Z","2014-06-28T10:00:28Z","2012-05"],"dc:description":["Made available in DSpace on 2012-06-27T21:31:00Z (GMT). No. of bitstreams: 2 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5) license.txt: 4066 bytes, checksum: e11c5bd6828b3d05964feec70e3062c8 (MD5)","Item marked as restricted to the 'Administrator' Group (id=1) by William Ingram (wingram2@illinois.edu) on 2012-06-27T21:32:46Z Item is restricted until 2014-06-27T21:32:23Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:28Z Item was in collections: Graduate Theses and Dissertations at Illinois (ID: 204) Dissertations and Theses - Computer Science (ID: 587) No. of bitstreams: 2 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5) license.txt: 4066 bytes, checksum: e11c5bd6828b3d05964feec70e3062c8 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2014-06-28T10:00:28Z","In this thesis I introduce visual phrases, complex visual composites like ``a person riding a horse''. Visual phrases often display significantly reduced visual complexity compared to their component objects, because the appearance of those objects can change profoundly when they participate in relations. I introduce a dataset suitable for phrasal recognition that uses familiar PASCAL object categories, and demonstrate significant experimental gains resulting from exploiting visual phrases. I show that a visual phrase detector significantly outperforms a baseline which detects component objects and reasons about relations, even though visual phrase training sets tend to be smaller than those for objects. I argue that any multi-class detection system must decode detector outputs to produce final results; this is usually done with non-maximum suppression. I describe a novel decoding procedure that can account accurately for local context without solving difficult inference problems. I show this decoding procedure outperforms the state of the art. Finally, I show that decoding a combination of phrasal and object detectors produces real improvements in detector results.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2012-04-25T20:23:32Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 1 Sadeghi_Mohammad_Amin.pdf: 12334753 bytes, checksum: 738a2c76fd1c4928e84a9747bcd9b451 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/32060"],"dc:language":["en"],"dc:rights":["Copyright 2012 Mohammad Amin Sadeghi and Ali Farhadi under Creative Commons"],"dc:subject":["Visual Phrase","Phrasal Recognition","Visual Composites","Object Recognition"],"dc:title":["Recognition using visual phrases"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:30Z"}