{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/117682"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/117682","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Study on speech emotion recognition based on deep learning","abstract":"Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-12-01","abstract_html":"Submission published under a 24 month embargo labeled &#x27;U of I Access&#x27;, the embargo will last until 2024-12-01","abstract_has_math":false,"creators":["Guan, Haozhong"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-12","date_published":"2022-12","updated_at":"2026-07-22T22:24:56Z","subjects":["Speech Emotion Recognition","Convolution Neural Network","Resnet50"],"languages":["en","eng"],"rights":["Copyright 2022 Haozhong Guan"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/117682","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark"]},{"key":"dc:creator","label":"Author","values":["Guan, Haozhong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-12","2022-12-05"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Speech Emotion Recognition","Convolution Neural Network","Resnet50"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en","eng"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2022 Haozhong Guan"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/117682"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-12-01","The student, Haozhong Guan, accepted the attached license on 2022-12-02 at 16:04.","The student, Haozhong Guan, submitted this Thesis for approval on 2022-12-02 at 16:11.","This Thesis was approved for publication on 2022-12-05 at 14:06.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18723 on 2023-04-12 at 08:13:39","Speech emotion recognition (SER) is closely related to human life, and has the potential to bring great changes and improvements to people's lives. The continuous development of artificial intelligence and SER will bring new breakthroughs to the field of human-machine interaction. Therefore, studying SER has extremely important theoretical value and research significance. In this thesis, the development status of speech emotion recognition is reviewed, and the existing problems and development challenges are pointed out. On the basis of summarizing the key technologies of speech emotion recognition, the speech emotion recognition model of ResNet50 CNN is constructed, and the recognition experiment and analysis are carried out. The main work is as follows: The speech emotion description model, the process of speech emotion recognition, the preprocessing of speech signals and the extraction method of emotion feature parameters are summarized. The time domain waveform and the spectrogram characteristics of different emotional speeches are analyzed, and the speech emotion recognition scheme combining the extraction of spectrogram and CNN is determined. In this thesis, a CNN model is constructed based on a residual network, which uses ResNet50 network and bottleneck block, and consists of 49 convolutional layers and one fully connected layer. The output is expressed as a linear superposition of nonlinear transformation by “shortcut connections” of residual network, which improves the problem of gradient disappearance or explosion in the process of back propagation, and makes the deep network get better training. Based on IEMOCAP and Emo-DB datasets, the efficient speech emotion recognition is realized. The results show that the recognition accuracies of the constructed ResNet50 CNN model for IEMOCAP and Emo-DB datasets are 69.12% and 85.92%, respectively. Compared with other deep learning models, the proposed ResNet50 CNN model is simple and efficient."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Study on speech emotion recognition based on deep learning"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark"],"dc:creator":["Guan, Haozhong"],"dc:date":["2022-12","2022-12-05"],"dc:description":["Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2024-12-01","The student, Haozhong Guan, accepted the attached license on 2022-12-02 at 16:04.","The student, Haozhong Guan, submitted this Thesis for approval on 2022-12-02 at 16:11.","This Thesis was approved for publication on 2022-12-05 at 14:06.","DSpace SAF Submission Ingestion Package generated from Vireo submission #18723 on 2023-04-12 at 08:13:39","Speech emotion recognition (SER) is closely related to human life, and has the potential to bring great changes and improvements to people's lives. The continuous development of artificial intelligence and SER will bring new breakthroughs to the field of human-machine interaction. Therefore, studying SER has extremely important theoretical value and research significance. In this thesis, the development status of speech emotion recognition is reviewed, and the existing problems and development challenges are pointed out. On the basis of summarizing the key technologies of speech emotion recognition, the speech emotion recognition model of ResNet50 CNN is constructed, and the recognition experiment and analysis are carried out. The main work is as follows: The speech emotion description model, the process of speech emotion recognition, the preprocessing of speech signals and the extraction method of emotion feature parameters are summarized. The time domain waveform and the spectrogram characteristics of different emotional speeches are analyzed, and the speech emotion recognition scheme combining the extraction of spectrogram and CNN is determined. In this thesis, a CNN model is constructed based on a residual network, which uses ResNet50 network and bottleneck block, and consists of 49 convolutional layers and one fully connected layer. The output is expressed as a linear superposition of nonlinear transformation by “shortcut connections” of residual network, which improves the problem of gradient disappearance or explosion in the process of back propagation, and makes the deep network get better training. Based on IEMOCAP and Emo-DB datasets, the efficient speech emotion recognition is realized. The results show that the recognition accuracies of the constructed ResNet50 CNN model for IEMOCAP and Emo-DB datasets are 69.12% and 85.92%, respectively. Compared with other deep learning models, the proposed ResNet50 CNN model is simple and efficient."],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/117682"],"dc:language":["en","eng"],"dc:rights":["Copyright 2022 Haozhong Guan"],"dc:subject":["Speech Emotion Recognition","Convolution Neural Network","Resnet50"],"dc:title":["Study on speech emotion recognition based on deep learning"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:56Z"}