{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/108579"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/108579","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"A journey to photo-realistic facial animation synthesis","abstract":"This dissertation presents preliminary work in facial animation generation with applications in educational psychology. In the ﬁrst part of this dissertation, we describe two psychology studies as well as the computer vision techniques and platforms being used. Both studies investigate using conversational agents (CAs) as a way of delivering medical messages to patients. By incorporating CAs in the system, both semantic and emotional information can be delivered, which helps the patients, especially those with low heath and numerical literacy, to get a better understanding of their test results and medical instructions. Human studies were conducted to test the eﬀectiveness of CA. In addition, the whole system was integrated with speech recognition as well as a natural language processing module to enable teach-back capability of the CA by providing a correct answer in case of a given wrong answer by the user, which will help them to get a better understanding of the medical messages being delivered. The second part of this dissertation documents the details of a proposed neural network based facial animation synthesis method. By unifying both appearance-based and warping-based methods in an end-to-end training process, the proposed system was able to generate vivid facial animation with highly preserved details. We show both qualitatively and quantitatively that the proposed system achieved a higher performance than baseline methods. In addition, visualization and ablation studies were conducted to further justify the eﬀectiveness of the proposed system. In the third part, the previous facial animation synthesis system was integrated with another audio speech processing system. The ﬁnal system was able to take speech signal and sample face images as input and generate the corresponding talking head animation as output. Comparison to the previous state-of-the-art method shows that the proposed system in this work achieves better performance.","abstract_html":"This dissertation presents preliminary work in facial animation generation with applications in educational psychology. In the ﬁrst part of this dissertation, we describe two psychology studies as well as the computer vision techniques and platforms being used. Both studies investigate using conversational agents (CAs) as a way of delivering medical messages to patients. By incorporating CAs in the system, both semantic and emotional information can be delivered, which helps the patients, especially those with low heath and numerical literacy, to get a better understanding of their test results and medical instructions. Human studies were conducted to test the eﬀectiveness of CA. In addition, the whole system was integrated with speech recognition as well as a natural language processing module to enable teach-back capability of the CA by providing a correct answer in case of a given wrong answer by the user, which will help them to get a better understanding of the medical messages being delivered. The second part of this dissertation documents the details of a proposed neural network based facial animation synthesis method. By unifying both appearance-based and warping-based methods in an end-to-end training process, the proposed system was able to generate vivid facial animation with highly preserved details. We show both qualitatively and quantitatively that the proposed system achieved a higher performance than baseline methods. In addition, visualization and ablation studies were conducted to further justify the eﬀectiveness of the proposed system. In the third part, the previous facial animation synthesis system was integrated with another audio speech processing system. The ﬁnal system was able to take speech signal and sample face images as input and generate the corresponding talking head animation as output. Comparison to the previous state-of-the-art method shows that the proposed system in this work achieves better performance.","abstract_has_math":false,"creators":["Gu, Kuangxiao"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Electrical & Computer Engr","degree_department":null,"school":null,"contributors":["Hasegawa-Johnson, Mark","Huang, Thomas S.","Morrow, Daniel G.","Liang, Zhi-Pei","Shi, Honghui"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-10-07T22:44:24Z","date_published":"2020-10-07T22:44:24Z","updated_at":"2026-07-22T22:24:48Z","subjects":["facial animation synthesis","talking head","patient portal","deep learning","neural network","audio-driven facial animation synthesis"],"languages":["en"],"rights":["Copyright 2020 Kuangxiao Gu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/108579","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Hasegawa-Johnson, Mark","Huang, Thomas S.","Morrow, Daniel G.","Liang, Zhi-Pei","Shi, Honghui"]},{"key":"dc:creator","label":"Author","values":["Gu, Kuangxiao"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-10-07T22:44:24Z","2022-10-07T22:44:53Z","2020-07-08","2020-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical & Computer Engr"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["facial animation synthesis","talking head","patient portal","deep learning","neural network","audio-driven facial animation synthesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2020 Kuangxiao Gu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/108579"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This dissertation presents preliminary work in facial animation generation with applications in educational psychology. In the ﬁrst part of this dissertation, we describe two psychology studies as well as the computer vision techniques and platforms being used. Both studies investigate using conversational agents (CAs) as a way of delivering medical messages to patients. By incorporating CAs in the system, both semantic and emotional information can be delivered, which helps the patients, especially those with low heath and numerical literacy, to get a better understanding of their test results and medical instructions. Human studies were conducted to test the eﬀectiveness of CA. In addition, the whole system was integrated with speech recognition as well as a natural language processing module to enable teach-back capability of the CA by providing a correct answer in case of a given wrong answer by the user, which will help them to get a better understanding of the medical messages being delivered. The second part of this dissertation documents the details of a proposed neural network based facial animation synthesis method. By unifying both appearance-based and warping-based methods in an end-to-end training process, the proposed system was able to generate vivid facial animation with highly preserved details. We show both qualitatively and quantitatively that the proposed system achieved a higher performance than baseline methods. In addition, visualization and ablation studies were conducted to further justify the eﬀectiveness of the proposed system. In the third part, the previous facial animation synthesis system was integrated with another audio speech processing system. The ﬁnal system was able to take speech signal and sample face images as input and generate the corresponding talking head animation as output. Comparison to the previous state-of-the-art method shows that the proposed system in this work achieves better performance.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-08-01","The student, Kuangxiao Gu, accepted the attached license on 2020-07-07 at 09:57.","The student, Kuangxiao Gu, submitted this Dissertation for approval on 2020-07-07 at 10:02.","This Dissertation was approved for publication on 2020-07-08 at 09:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15508 on 2020-10-02 at 15:31:22","Made available in DSpace on 2020-10-07T22:44:24Z (GMT). No. of bitstreams: 3 GU-DISSERTATION-2020.pdf: 26282577 bytes, checksum: 355d4d9b6204bf75f71be7de1ce138ae (MD5) LICENSE.txt: 4209 bytes, checksum: 21eb7a60c3c1d2974df1aaf96acbad63 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: b59ef3efc5459fc3c6f120e9b96ddfc1 (MD5) Previous issue date: 2020-07-08","Embargo set by: Seth Robbins for item 116206 Lift date: 2022-10-07T22:44:53Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["A journey to photo-realistic facial animation synthesis"]}]}],"canonical_facts":{"dc:contributor":["Hasegawa-Johnson, Mark","Huang, Thomas S.","Morrow, Daniel G.","Liang, Zhi-Pei","Shi, Honghui"],"dc:creator":["Gu, Kuangxiao"],"dc:date":["2020-10-07T22:44:24Z","2022-10-07T22:44:53Z","2020-07-08","2020-08"],"dc:description":["This dissertation presents preliminary work in facial animation generation with applications in educational psychology. In the ﬁrst part of this dissertation, we describe two psychology studies as well as the computer vision techniques and platforms being used. Both studies investigate using conversational agents (CAs) as a way of delivering medical messages to patients. By incorporating CAs in the system, both semantic and emotional information can be delivered, which helps the patients, especially those with low heath and numerical literacy, to get a better understanding of their test results and medical instructions. Human studies were conducted to test the eﬀectiveness of CA. In addition, the whole system was integrated with speech recognition as well as a natural language processing module to enable teach-back capability of the CA by providing a correct answer in case of a given wrong answer by the user, which will help them to get a better understanding of the medical messages being delivered. The second part of this dissertation documents the details of a proposed neural network based facial animation synthesis method. By unifying both appearance-based and warping-based methods in an end-to-end training process, the proposed system was able to generate vivid facial animation with highly preserved details. We show both qualitatively and quantitatively that the proposed system achieved a higher performance than baseline methods. In addition, visualization and ablation studies were conducted to further justify the eﬀectiveness of the proposed system. In the third part, the previous facial animation synthesis system was integrated with another audio speech processing system. The ﬁnal system was able to take speech signal and sample face images as input and generate the corresponding talking head animation as output. Comparison to the previous state-of-the-art method shows that the proposed system in this work achieves better performance.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2022-08-01","The student, Kuangxiao Gu, accepted the attached license on 2020-07-07 at 09:57.","The student, Kuangxiao Gu, submitted this Dissertation for approval on 2020-07-07 at 10:02.","This Dissertation was approved for publication on 2020-07-08 at 09:53.","DSpace SAF Submission Ingestion Package generated from Vireo submission #15508 on 2020-10-02 at 15:31:22","Made available in DSpace on 2020-10-07T22:44:24Z (GMT). No. of bitstreams: 3 GU-DISSERTATION-2020.pdf: 26282577 bytes, checksum: 355d4d9b6204bf75f71be7de1ce138ae (MD5) LICENSE.txt: 4209 bytes, checksum: 21eb7a60c3c1d2974df1aaf96acbad63 (MD5) PROQUEST_LICENSE.txt: 4555 bytes, checksum: b59ef3efc5459fc3c6f120e9b96ddfc1 (MD5) Previous issue date: 2020-07-08","Embargo set by: Seth Robbins for item 116206 Lift date: 2022-10-07T22:44:53Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/108579"],"dc:language":["en"],"dc:rights":["Copyright 2020 Kuangxiao Gu"],"dc:subject":["facial animation synthesis","talking head","patient portal","deep learning","neural network","audio-driven facial animation synthesis"],"dc:title":["A journey to photo-realistic facial animation synthesis"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Electrical & Computer Engr"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:48Z"}