{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/81332"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/81332","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Joint Processing of Audio -Visual Information for the Recognition of Emotional Expressions in Human -Computer Interaction","abstract":"\"This thesis addresses the problem of detecting human emotional expressions by computer from the voice and facial motions of the user. The computer is equipped with a microphone to listen to the user's voice, and a video camera to look at the user. Prosodic features in the audio and facial motions exhibited on the face can help the computer make some inferences about the user's emotional state, assuming the users are willing to show their emotions. Another problem it addresses is the coupling between voice and the facial expression. Sometimes the user moves the lips to produce the speech, and sometimes the user only exhibits facial expression without speaking any words. Therefore, it is important to handle these two modalities accordingly. In particular, a pure \"\"facial expression detector\"\" will not function properly when the person is speaking, and a pure \"\"vocal emotion recognizer\"\" is useless when the user is not speaking. In this thesis, a complementary relationship between audio and video is proposed. Although these two modalities do not couple strongly in time, they seem to complement each other. In some cases, similar facial expressions may have different vocal characteristics, and vocal emotions having similar properties may have distinct facial behaviors.\"","abstract_html":"&quot;This thesis addresses the problem of detecting human emotional expressions by computer from the voice and facial motions of the user. The computer is equipped with a microphone to listen to the user&#x27;s voice, and a video camera to look at the user. Prosodic features in the audio and facial motions exhibited on the face can help the computer make some inferences about the user&#x27;s emotional state, assuming the users are willing to show their emotions. Another problem it addresses is the coupling between voice and the facial expression. Sometimes the user moves the lips to produce the speech, and sometimes the user only exhibits facial expression without speaking any words. Therefore, it is important to handle these two modalities accordingly. In particular, a pure &quot;&quot;facial expression detector&quot;&quot; will not function properly when the person is speaking, and a pure &quot;&quot;vocal emotion recognizer&quot;&quot; is useless when the user is not speaking. In this thesis, a complementary relationship between audio and video is proposed. Although these two modalities do not couple strongly in time, they seem to complement each other. In some cases, similar facial expressions may have different vocal characteristics, and vocal emotions having similar properties may have distinct facial behaviors.&quot;","abstract_has_math":false,"creators":["Chen, Lawrence Shao-Hsien"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical Engineering","degree_department":null,"school":null,"contributors":["Huang, Thomas S."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-09-25T20:10:37Z","date_published":"2015-09-25T20:10:37Z","updated_at":"2026-07-22T22:26:16Z","subjects":["Computer Science"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(MiAaPQ)AAI9971046"],"render_values":[{"text":"(MiAaPQ)AAI9971046","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/81332","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Thomas S."]},{"key":"dc:creator","label":"Author","values":["Chen, Lawrence Shao-Hsien"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2015-09-25T20:10:37Z","10000-01-01","2000"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/81332","(MiAaPQ)AAI9971046"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"This thesis addresses the problem of detecting human emotional expressions by computer from the voice and facial motions of the user. The computer is equipped with a microphone to listen to the user's voice, and a video camera to look at the user. Prosodic features in the audio and facial motions exhibited on the face can help the computer make some inferences about the user's emotional state, assuming the users are willing to show their emotions. Another problem it addresses is the coupling between voice and the facial expression. Sometimes the user moves the lips to produce the speech, and sometimes the user only exhibits facial expression without speaking any words. Therefore, it is important to handle these two modalities accordingly. In particular, a pure \"\"facial expression detector\"\" will not function properly when the person is speaking, and a pure \"\"vocal emotion recognizer\"\" is useless when the user is not speaking. In this thesis, a complementary relationship between audio and video is proposed. Although these two modalities do not couple strongly in time, they seem to complement each other. In some cases, similar facial expressions may have different vocal characteristics, and vocal emotions having similar properties may have distinct facial behaviors.\"","Made available in DSpace on 2015-09-25T20:10:37Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 9971046.pdf: 3585097 bytes, checksum: 794117d90d3d5fdeb3a32091c9a4a802 (MD5) Previous issue date: 2000","Embargo set by: Seth Robbins for item 82613 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","66 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2000."]},{"key":"dc:title","label":"Title","values":["Joint Processing of Audio -Visual Information for the Recognition of Emotional Expressions in Human -Computer Interaction"]}]}],"canonical_facts":{"dc:contributor":["Huang, Thomas S."],"dc:creator":["Chen, Lawrence Shao-Hsien"],"dc:date":["2015-09-25T20:10:37Z","10000-01-01","2000"],"dc:description":["\"This thesis addresses the problem of detecting human emotional expressions by computer from the voice and facial motions of the user. The computer is equipped with a microphone to listen to the user's voice, and a video camera to look at the user. Prosodic features in the audio and facial motions exhibited on the face can help the computer make some inferences about the user's emotional state, assuming the users are willing to show their emotions. Another problem it addresses is the coupling between voice and the facial expression. Sometimes the user moves the lips to produce the speech, and sometimes the user only exhibits facial expression without speaking any words. Therefore, it is important to handle these two modalities accordingly. In particular, a pure \"\"facial expression detector\"\" will not function properly when the person is speaking, and a pure \"\"vocal emotion recognizer\"\" is useless when the user is not speaking. In this thesis, a complementary relationship between audio and video is proposed. Although these two modalities do not couple strongly in time, they seem to complement each other. In some cases, similar facial expressions may have different vocal characteristics, and vocal emotions having similar properties may have distinct facial behaviors.\"","Made available in DSpace on 2015-09-25T20:10:37Z (GMT). No. of bitstreams: 2 license.txt: 4848 bytes, checksum: 96035ab3f5e1c23cc7138a224ce498bd (MD5) 9971046.pdf: 3585097 bytes, checksum: 794117d90d3d5fdeb3a32091c9a4a802 (MD5) Previous issue date: 2000","Embargo set by: Seth Robbins for item 82613 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","66 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 2000."],"dc:identifier":["http://hdl.handle.net/2142/81332","(MiAaPQ)AAI9971046"],"dc:language":["eng"],"dc:subject":["Computer Science"],"dc:title":["Joint Processing of Audio -Visual Information for the Recognition of Emotional Expressions in Human -Computer Interaction"],"dc:type":["text"],"thesis:degree_discipline":["Electrical Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:16Z"}