{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/89047"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/89047","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Fast object detection","abstract":"The ultimate goal of computer vision is to understand images. We describe methods to understand images at two levels. One is at the level of description of images which we produce using sentences. These sentences talk about the things that are present in the image and about where they are and what they are doing. Then we ask in what ways should we describe images. We introduce visual phrases that are composite chunks of meaning. We show that object detectors could be better at detecting some visual phrases than detecting single objects. This process of image understanding needs to use a lot of detectors. Running conventional object detectors at the rate required for image understanding could be very slow. We study fast object detection from an engineering perspective. We argue that a desirable object detector must: (1) be able to work with legacy templates; (2) be random access; (3) be able to trade accuracy versus speed; (4) have any-time property. We describe a method to have all of these features together for a fast detector. We apply these techniques to deformable parts model object detectors and show two orders of magnitude speed-up while adding their desirable features. We finally investigate the consequences of this architecture with a view of improving convolutional neural networks.","abstract_html":"The ultimate goal of computer vision is to understand images. We describe methods to understand images at two levels. One is at the level of description of images which we produce using sentences. These sentences talk about the things that are present in the image and about where they are and what they are doing. Then we ask in what ways should we describe images. We introduce visual phrases that are composite chunks of meaning. We show that object detectors could be better at detecting some visual phrases than detecting single objects. This process of image understanding needs to use a lot of detectors. Running conventional object detectors at the rate required for image understanding could be very slow. We study fast object detection from an engineering perspective. We argue that a desirable object detector must: (1) be able to work with legacy templates; (2) be random access; (3) be able to trade accuracy versus speed; (4) have any-time property. We describe a method to have all of these features together for a fast detector. We apply these techniques to deformable parts model object detectors and show two orders of magnitude speed-up while adding their desirable features. We finally investigate the consequences of this architecture with a view of improving convolutional neural networks.","abstract_has_math":false,"creators":["Sadeghi, Mohammad Amin"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Forsyth, David A","Hoiem, Derek","Ramanan, Deva","Lazebnik, Svetlana","Golparvar-fard, Mani"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016-03-02T19:34:19Z","date_published":"2016-03-02T19:34:19Z","updated_at":"2026-07-22T22:26:32Z","subjects":["Fast Object Detection","Accurate Object Detection","Visual Phrases","Sentence Generation","Vector Quantization."],"languages":["en"],"rights":["Copyright 2015 Mohammad Amin Sadeghi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/89047","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Forsyth, David A","Hoiem, Derek","Ramanan, Deva","Lazebnik, Svetlana","Golparvar-fard, Mani"]},{"key":"dc:creator","label":"Author","values":["Sadeghi, Mohammad Amin"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2016-03-02T19:34:19Z","2015-12-04","2015-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Fast Object Detection","Accurate Object Detection","Visual Phrases","Sentence Generation","Vector Quantization."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2015 Mohammad Amin Sadeghi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/89047"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The ultimate goal of computer vision is to understand images. We describe methods to understand images at two levels. One is at the level of description of images which we produce using sentences. These sentences talk about the things that are present in the image and about where they are and what they are doing. Then we ask in what ways should we describe images. We introduce visual phrases that are composite chunks of meaning. We show that object detectors could be better at detecting some visual phrases than detecting single objects. This process of image understanding needs to use a lot of detectors. Running conventional object detectors at the rate required for image understanding could be very slow. We study fast object detection from an engineering perspective. We argue that a desirable object detector must: (1) be able to work with legacy templates; (2) be random access; (3) be able to trade accuracy versus speed; (4) have any-time property. We describe a method to have all of these features together for a fast detector. We apply these techniques to deformable parts model object detectors and show two orders of magnitude speed-up while adding their desirable features. We finally investigate the consequences of this architecture with a view of improving convolutional neural networks.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-03-02 without embargo terms","The student, Mohammad Amin Sadeghi, accepted the attached license on 2015-12-03 at 19:25.","The student, Mohammad Amin Sadeghi, submitted this Dissertation for approval on 2015-12-03 at 19:42.","This Dissertation was approved for publication on 2015-12-04 at 15:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8930 on 2016-03-02 at 12:51:24","Made available in DSpace on 2016-03-02T19:34:19Z (GMT). No. of bitstreams: 2 SADEGHI-DISSERTATION-2015.pdf: 17726072 bytes, checksum: a9926b4b93154ebbd4c6cd886df92dbe (MD5) LICENSE.txt: 4218 bytes, checksum: a106486153c1ec5b3e0207c6ec0c9232 (MD5) Previous issue date: 2015-12-04"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Fast object detection"]}]}],"canonical_facts":{"dc:contributor":["Forsyth, David A","Hoiem, Derek","Ramanan, Deva","Lazebnik, Svetlana","Golparvar-fard, Mani"],"dc:creator":["Sadeghi, Mohammad Amin"],"dc:date":["2016-03-02T19:34:19Z","2015-12-04","2015-12"],"dc:description":["The ultimate goal of computer vision is to understand images. We describe methods to understand images at two levels. One is at the level of description of images which we produce using sentences. These sentences talk about the things that are present in the image and about where they are and what they are doing. Then we ask in what ways should we describe images. We introduce visual phrases that are composite chunks of meaning. We show that object detectors could be better at detecting some visual phrases than detecting single objects. This process of image understanding needs to use a lot of detectors. Running conventional object detectors at the rate required for image understanding could be very slow. We study fast object detection from an engineering perspective. We argue that a desirable object detector must: (1) be able to work with legacy templates; (2) be random access; (3) be able to trade accuracy versus speed; (4) have any-time property. We describe a method to have all of these features together for a fast detector. We apply these techniques to deformable parts model object detectors and show two orders of magnitude speed-up while adding their desirable features. We finally investigate the consequences of this architecture with a view of improving convolutional neural networks.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2016-03-02 without embargo terms","The student, Mohammad Amin Sadeghi, accepted the attached license on 2015-12-03 at 19:25.","The student, Mohammad Amin Sadeghi, submitted this Dissertation for approval on 2015-12-03 at 19:42.","This Dissertation was approved for publication on 2015-12-04 at 15:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #8930 on 2016-03-02 at 12:51:24","Made available in DSpace on 2016-03-02T19:34:19Z (GMT). No. of bitstreams: 2 SADEGHI-DISSERTATION-2015.pdf: 17726072 bytes, checksum: a9926b4b93154ebbd4c6cd886df92dbe (MD5) LICENSE.txt: 4218 bytes, checksum: a106486153c1ec5b3e0207c6ec0c9232 (MD5) Previous issue date: 2015-12-04"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/89047"],"dc:language":["en"],"dc:rights":["Copyright 2015 Mohammad Amin Sadeghi"],"dc:subject":["Fast Object Detection","Accurate Object Detection","Visual Phrases","Sentence Generation","Vector Quantization."],"dc:title":["Fast object detection"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:32Z"}