{"id":{"repo_id":"calpoly","oai_identifier":"oai:digitalcommons.calpoly.edu:theses-2586"},"canonical_url":"https://search.dev.ndltd.org/etd/calpoly/oai:digitalcommons.calpoly.edu:theses-2586","repository":{"repo_id":"calpoly","name":"Cal Poly","base_url":"https://digitalcommons.calpoly.edu/do/oai/"},"display":{"title":"IR-Depth Face Detection and Lip Localization Using Kinect V2","abstract":"<p>Face recognition and lip localization are two main building blocks in the development of audio visual automatic speech recognition systems (AV-ASR). In many earlier works, face recognition and lip localization were conducted in uniform lighting conditions with simple backgrounds. However, such conditions are seldom the case in real world applications. In this paper, we present an approach to face recognition and lip localization that is invariant to lighting conditions. This is done by employing infrared and depth images captured by the Kinect V2 device. First we present the use of infrared images for face detection. Second, we use the face’s inherent depth information to reduce the search area for the lips by developing a nose point detection. Third, we further reduce the search area by using a depth segmentation algorithm to separate the face from its background. Finally, with the reduced search range, we present a method for lip localization based on depth gradients. Experimental results demonstrated an accuracy of 100% for face detection, and 96% for lip localization.</p>","abstract_html":"&lt;p&gt;Face recognition and lip localization are two main building blocks in the development of audio visual automatic speech recognition systems (AV-ASR). In many earlier works, face recognition and lip localization were conducted in uniform lighting conditions with simple backgrounds. However, such conditions are seldom the case in real world applications. In this paper, we present an approach to face recognition and lip localization that is invariant to lighting conditions. This is done by employing infrared and depth images captured by the Kinect V2 device. First we present the use of infrared images for face detection. Second, we use the face’s inherent depth information to reduce the search area for the lips by developing a nose point detection. Third, we further reduce the search area by using a depth segmentation algorithm to separate the face from its background. Finally, with the reduced search range, we present a method for lip localization based on depth gradients. Experimental results demonstrated an accuracy of 100% for face detection, and 96% for lip localization.&lt;/p&gt;","abstract_has_math":false,"creators":["Fong, Katherine Kayan"],"institution":null,"degree_name":"MS in Electrical Engineering","degree_level":null,"degree_discipline":"Electrical Engineering","degree_department":null,"school":null,"contributors":["Xiaozheng (Jane) Zhang"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-06-01T07:00:00Z","date_published":"2015-06-01T07:00:00Z","updated_at":"2026-07-24T01:32:37Z","subjects":["Face detection","lip localization","infrared (IR)","Audio-visual automatic speech recognition","depth information","Microsoft Kinect","Signal Processing"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["10.15368/theses.2015.81"],"render_values":[{"text":"10.15368/theses.2015.81","href":"https://doi.org/10.15368/theses.2015.81","code":true}]}]},"links":{"outbound_url":"https://digitalcommons.calpoly.edu/theses/1425","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Xiaozheng (Jane) Zhang"]},{"key":"dc:creator","label":"Author","values":["Fong, Katherine Kayan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2015-06-19T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS in Electrical Engineering"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Face detection","lip localization","infrared (IR)","Audio-visual automatic speech recognition","depth information","Microsoft Kinect","Signal Processing"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.calpoly.edu/theses/1425","10.15368/theses.2015.81"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Face recognition and lip localization are two main building blocks in the development of audio visual automatic speech recognition systems (AV-ASR). In many earlier works, face recognition and lip localization were conducted in uniform lighting conditions with simple backgrounds. However, such conditions are seldom the case in real world applications. In this paper, we present an approach to face recognition and lip localization that is invariant to lighting conditions. This is done by employing infrared and depth images captured by the Kinect V2 device. First we present the use of infrared images for face detection. Second, we use the face’s inherent depth information to reduce the search area for the lips by developing a nose point detection. Third, we further reduce the search area by using a depth segmentation algorithm to separate the face from its background. Finally, with the reduced search range, we present a method for lip localization based on depth gradients. Experimental results demonstrated an accuracy of 100% for face detection, and 96% for lip localization.</p>"]},{"key":"dc:title","label":"Title","values":["IR-Depth Face Detection and Lip Localization Using Kinect V2"]}]}],"canonical_facts":{"dc:contributor":["Xiaozheng (Jane) Zhang"],"dc:creator":["Fong, Katherine Kayan"],"dc:date.available":["2015-06-19T07:00:00Z"],"dc:description.abstract":["<p>Face recognition and lip localization are two main building blocks in the development of audio visual automatic speech recognition systems (AV-ASR). In many earlier works, face recognition and lip localization were conducted in uniform lighting conditions with simple backgrounds. However, such conditions are seldom the case in real world applications. In this paper, we present an approach to face recognition and lip localization that is invariant to lighting conditions. This is done by employing infrared and depth images captured by the Kinect V2 device. First we present the use of infrared images for face detection. Second, we use the face’s inherent depth information to reduce the search area for the lips by developing a nose point detection. Third, we further reduce the search area by using a depth segmentation algorithm to separate the face from its background. Finally, with the reduced search range, we present a method for lip localization based on depth gradients. Experimental results demonstrated an accuracy of 100% for face detection, and 96% for lip localization.</p>"],"dc:identifier":["https://digitalcommons.calpoly.edu/theses/1425","10.15368/theses.2015.81"],"dc:subject":["Face detection","lip localization","infrared (IR)","Audio-visual automatic speech recognition","depth information","Microsoft Kinect","Signal Processing"],"dc:title":["IR-Depth Face Detection and Lip Localization Using Kinect V2"],"thesis:degree_discipline":["Electrical Engineering"],"thesis:degree_name":["MS in Electrical Engineering"]},"updated_at":"2026-07-24T01:32:37Z"}