{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/150226"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/150226","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Restoring Eye Contact in Video Conferencing","abstract":"In recent years, especially during the COVID-19 pandemic, video conferencing applications have become widely adopted, enabling remote work and virtual learning. Despite the convenience, video conferencing made it challenging and even unnatural to establish eye contact, which is a critical component in visual communication. To create the perception of eye contact in a video conferencing call, a user would need to look directly into the camera, but the user typically looks at the participant displayed on the screen while the camera is located at the top of the screen. Such physical deviation results in the perception that the user is looking elsewhere to the other participant. This work proposes the application of a Convolutional Neural Network (CNN) based 3D face reconstruction technique, Position Map Regression Network (PRNet), on 2D images from a single RGB webcam found in consumer-grade computers to create a newly synthesized video stream where the video conferencing user’s face becomes oriented towards the webcam, resolving the physical deviation between the webcam location and the location of the other participant shown on the screen. Unlike previous approaches, this work fits a pre-trained model onto the specific user to leverage more accurate 3D face reconstruction results.","abstract_html":"In recent years, especially during the COVID-19 pandemic, video conferencing applications have become widely adopted, enabling remote work and virtual learning. Despite the convenience, video conferencing made it challenging and even unnatural to establish eye contact, which is a critical component in visual communication. To create the perception of eye contact in a video conferencing call, a user would need to look directly into the camera, but the user typically looks at the participant displayed on the screen while the camera is located at the top of the screen. Such physical deviation results in the perception that the user is looking elsewhere to the other participant. This work proposes the application of a Convolutional Neural Network (CNN) based 3D face reconstruction technique, Position Map Regression Network (PRNet), on 2D images from a single RGB webcam found in consumer-grade computers to create a newly synthesized video stream where the video conferencing user’s face becomes oriented towards the webcam, resolving the physical deviation between the webcam location and the location of the other participant shown on the screen. Unlike previous approaches, this work fits a pre-trained model onto the specific user to leverage more accurate 3D face reconstruction results.","abstract_has_math":false,"creators":["Kim, Jin Woo"],"institution":"Massachusetts Institute of Technology","degree_name":"Master","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science","school":null,"contributors":[],"advisors":["Lee, Chong U.","Lim, Jae S."],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-02","date_published":"2023-02","updated_at":"2026-07-22T22:21:48Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"rights_urls":["http://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/150226","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lee, Chong U.","Lim, Jae S."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"]},{"key":"dc:creator","label":"Author","values":["Kim, Jin Woo"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2023-03-31T14:40:58Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2023-03-31T14:40:58Z"]},{"key":"dc:date.issued","label":"Date","values":["2023-02"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master","Master of Engineering in Electrical Engineering and Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright MIT"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/150226"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In recent years, especially during the COVID-19 pandemic, video conferencing applications have become widely adopted, enabling remote work and virtual learning. Despite the convenience, video conferencing made it challenging and even unnatural to establish eye contact, which is a critical component in visual communication. To create the perception of eye contact in a video conferencing call, a user would need to look directly into the camera, but the user typically looks at the participant displayed on the screen while the camera is located at the top of the screen. Such physical deviation results in the perception that the user is looking elsewhere to the other participant. This work proposes the application of a Convolutional Neural Network (CNN) based 3D face reconstruction technique, Position Map Regression Network (PRNet), on 2D images from a single RGB webcam found in consumer-grade computers to create a newly synthesized video stream where the video conferencing user’s face becomes oriented towards the webcam, resolving the physical deviation between the webcam location and the location of the other participant shown on the screen. Unlike previous approaches, this work fits a pre-trained model onto the specific user to leverage more accurate 3D face reconstruction results."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.Eng."]},{"key":"dc:title","label":"Title","values":["Restoring Eye Contact in Video Conferencing"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lee, Chong U.","Lim, Jae S."],"dc:contributor.department":["Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science"],"dc:creator":["Kim, Jin Woo"],"dc:date.accessioned":["2023-03-31T14:40:58Z"],"dc:date.available":["2023-03-31T14:40:58Z"],"dc:date.issued":["2023-02"],"dc:description.abstract":["In recent years, especially during the COVID-19 pandemic, video conferencing applications have become widely adopted, enabling remote work and virtual learning. Despite the convenience, video conferencing made it challenging and even unnatural to establish eye contact, which is a critical component in visual communication. To create the perception of eye contact in a video conferencing call, a user would need to look directly into the camera, but the user typically looks at the participant displayed on the screen while the camera is located at the top of the screen. Such physical deviation results in the perception that the user is looking elsewhere to the other participant. This work proposes the application of a Convolutional Neural Network (CNN) based 3D face reconstruction technique, Position Map Regression Network (PRNet), on 2D images from a single RGB webcam found in consumer-grade computers to create a newly synthesized video stream where the video conferencing user’s face becomes oriented towards the webcam, resolving the physical deviation between the webcam location and the location of the other participant shown on the screen. Unlike previous approaches, this work fits a pre-trained model onto the specific user to leverage more accurate 3D face reconstruction results."],"dc:description.degree":["M.Eng."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/150226"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"dc:rights.uri":["http://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Restoring Eye Contact in Video Conferencing"],"dc:type":["Thesis"],"thesis:degree_name":["Master","Master of Engineering in Electrical Engineering and Computer Science"]},"updated_at":"2026-07-22T22:21:48Z"}