{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/97284"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/97284","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"How deep learning can help emotion recognition","abstract":"As technological systems become more and more advanced, the need for including the human during the interaction process has become more apparent. One simple way is to have the computer system understand and respond to the human's emotions. Previous works in emotion recognition have focused on improving performance by incorporating domain knowledge into the underlying system either through pre-specified rules or hand-crafted features. However, in the last few years, learned feature representations have experienced a resurgence mainly due to the success of deep neural networks. In this dissertation, we highlight how deep neural networks, when applied to emotion recognition, can learn representations that not only achieve superior accuracy to hand-crafted techniques, but also align with previous domain knowledge. Moreover, we show how these learned representations can generalize to different definitions of emotions and to different input modalities. The first part of this dissertation considers the task of categorical emotion recognition on images. We show how a convolutional neural network (CNN) that achieves state-of-the-art performance can also learn features that strongly correspond to Facial Action Units (FAUs). In the second part, we focus our attention on emotion recognition in video. We take the image-based CNN model and combine it with a recurrent neural network (RNN) in order to do dimensional emotion recognition. We also visualize the portions of the faces that most strongly affect the output prediction by using the gradient as a saliency map. Lastly, we explore the merit of doing multimodal emotion recognition by combining our model with other models trained on audio and physiological data.","abstract_html":"As technological systems become more and more advanced, the need for including the human during the interaction process has become more apparent. One simple way is to have the computer system understand and respond to the human&#x27;s emotions. Previous works in emotion recognition have focused on improving performance by incorporating domain knowledge into the underlying system either through pre-specified rules or hand-crafted features. However, in the last few years, learned feature representations have experienced a resurgence mainly due to the success of deep neural networks. In this dissertation, we highlight how deep neural networks, when applied to emotion recognition, can learn representations that not only achieve superior accuracy to hand-crafted techniques, but also align with previous domain knowledge. Moreover, we show how these learned representations can generalize to different definitions of emotions and to different input modalities. The first part of this dissertation considers the task of categorical emotion recognition on images. We show how a convolutional neural network (CNN) that achieves state-of-the-art performance can also learn features that strongly correspond to Facial Action Units (FAUs). In the second part, we focus our attention on emotion recognition in video. We take the image-based CNN model and combine it with a recurrent neural network (RNN) in order to do dimensional emotion recognition. We also visualize the portions of the faces that most strongly affect the output prediction by using the gradient as a saliency map. Lastly, we explore the merit of doing multimodal emotion recognition by combining our model with other models trained on audio and physiological data.","abstract_has_math":false,"creators":["Khorrami, Pooya Rezvani"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Huang, Thomas S.","Hasegawa-Johnson, Mark","Hoiem, Derek W.","Liang, Zhi-Pei"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-08-10T19:14:36Z","date_published":"2017-08-10T19:14:36Z","updated_at":"2026-07-22T22:24:32Z","subjects":["Emotion recognition","Deep learning","Machine learning","Computer vision","Facial expression recognition","Affective computing","Deep neural networks"],"languages":["en"],"rights":["Copyright 2017 Pooya Rezvani Khorrami"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/97284","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Huang, Thomas S.","Hasegawa-Johnson, Mark","Hoiem, Derek W.","Liang, Zhi-Pei"]},{"key":"dc:creator","label":"Author","values":["Khorrami, Pooya Rezvani"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-08-10T19:14:36Z","2017-03-29","2017-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Emotion recognition","Deep learning","Machine learning","Computer vision","Facial expression recognition","Affective computing","Deep neural networks"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2017 Pooya Rezvani Khorrami"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/97284"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["As technological systems become more and more advanced, the need for including the human during the interaction process has become more apparent. One simple way is to have the computer system understand and respond to the human's emotions. Previous works in emotion recognition have focused on improving performance by incorporating domain knowledge into the underlying system either through pre-specified rules or hand-crafted features. However, in the last few years, learned feature representations have experienced a resurgence mainly due to the success of deep neural networks. In this dissertation, we highlight how deep neural networks, when applied to emotion recognition, can learn representations that not only achieve superior accuracy to hand-crafted techniques, but also align with previous domain knowledge. Moreover, we show how these learned representations can generalize to different definitions of emotions and to different input modalities. The first part of this dissertation considers the task of categorical emotion recognition on images. We show how a convolutional neural network (CNN) that achieves state-of-the-art performance can also learn features that strongly correspond to Facial Action Units (FAUs). In the second part, we focus our attention on emotion recognition in video. We take the image-based CNN model and combine it with a recurrent neural network (RNN) in order to do dimensional emotion recognition. We also visualize the portions of the faces that most strongly affect the output prediction by using the gradient as a saliency map. Lastly, we explore the merit of doing multimodal emotion recognition by combining our model with other models trained on audio and physiological data.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-08-10 without embargo terms","The student, Pooya Khorrami, accepted the attached license on 2017-03-22 at 15:42.","The student, Pooya Khorrami, submitted this Dissertation for approval on 2017-03-22 at 16:05.","This Dissertation was approved for publication on 2017-03-29 at 08:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10611 on 2017-08-10 at 13:38:18","Made available in DSpace on 2017-08-10T19:14:36Z (GMT). No. of bitstreams: 2 KHORRAMI-DISSERTATION-2017.pdf: 6929256 bytes, checksum: aff980313892634d5c7912135b0ce673 (MD5) LICENSE.txt: 4211 bytes, checksum: 3acfa05e953099228f390d2e409d01f9 (MD5) Previous issue date: 2017-03-29"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["How deep learning can help emotion recognition"]}]}],"canonical_facts":{"dc:contributor":["Huang, Thomas S.","Hasegawa-Johnson, Mark","Hoiem, Derek W.","Liang, Zhi-Pei"],"dc:creator":["Khorrami, Pooya Rezvani"],"dc:date":["2017-08-10T19:14:36Z","2017-03-29","2017-05"],"dc:description":["As technological systems become more and more advanced, the need for including the human during the interaction process has become more apparent. One simple way is to have the computer system understand and respond to the human's emotions. Previous works in emotion recognition have focused on improving performance by incorporating domain knowledge into the underlying system either through pre-specified rules or hand-crafted features. However, in the last few years, learned feature representations have experienced a resurgence mainly due to the success of deep neural networks. In this dissertation, we highlight how deep neural networks, when applied to emotion recognition, can learn representations that not only achieve superior accuracy to hand-crafted techniques, but also align with previous domain knowledge. Moreover, we show how these learned representations can generalize to different definitions of emotions and to different input modalities. The first part of this dissertation considers the task of categorical emotion recognition on images. We show how a convolutional neural network (CNN) that achieves state-of-the-art performance can also learn features that strongly correspond to Facial Action Units (FAUs). In the second part, we focus our attention on emotion recognition in video. We take the image-based CNN model and combine it with a recurrent neural network (RNN) in order to do dimensional emotion recognition. We also visualize the portions of the faces that most strongly affect the output prediction by using the gradient as a saliency map. Lastly, we explore the merit of doing multimodal emotion recognition by combining our model with other models trained on audio and physiological data.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-08-10 without embargo terms","The student, Pooya Khorrami, accepted the attached license on 2017-03-22 at 15:42.","The student, Pooya Khorrami, submitted this Dissertation for approval on 2017-03-22 at 16:05.","This Dissertation was approved for publication on 2017-03-29 at 08:54.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10611 on 2017-08-10 at 13:38:18","Made available in DSpace on 2017-08-10T19:14:36Z (GMT). No. of bitstreams: 2 KHORRAMI-DISSERTATION-2017.pdf: 6929256 bytes, checksum: aff980313892634d5c7912135b0ce673 (MD5) LICENSE.txt: 4211 bytes, checksum: 3acfa05e953099228f390d2e409d01f9 (MD5) Previous issue date: 2017-03-29"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/97284"],"dc:language":["en"],"dc:rights":["Copyright 2017 Pooya Rezvani Khorrami"],"dc:subject":["Emotion recognition","Deep learning","Machine learning","Computer vision","Facial expression recognition","Affective computing","Deep neural networks"],"dc:title":["How deep learning can help emotion recognition"],"dc:type":["text"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:32Z"}