{"id":{"repo_id":"carleton","oai_identifier":"oai:carleton.scholaris.ca:20.500.14718/42801"},"canonical_url":"https://search.dev.ndltd.org/etd/carleton/oai:carleton.scholaris.ca:20.500.14718/42801","repository":{"repo_id":"carleton","name":"Carleton University","base_url":"https://carleton.scholaris.ca/server/oai/request"},"display":{"title":"Deep Generative Models for Unsupervised Scale-Based and Position-Based Disentanglement of Concepts from Face Images.","abstract":"Among the different categories of natural images, face images are very important because of the role they play in human social interactions. It is recognised that despite all the recent advances of artificial intelligence using deep neural networks, computers are still struggling at achieving a rich and flexible understanding of face images comparable to humans&apos; face perception abilities. This thesis aims at finding fully unsupervised ways for learning a transformation from face images pixel space to a representation space in which the underlying facial concepts are captured and disentangled. We propose that it is possible to utilize clues from the real 3D world in order to guide the representation learner in the direction of disentangling facial concepts. We conduct two studies in order to test this hypothesis. First, we propose a deep autoencoder model for extracting facial concepts based on their scales. We introduce an adaptive resolution reconstruction loss inspired by the fact that different categories of concepts are encoded in (and can be captured from) different resolutions of face images. With the help of this new reconstruction loss, the deep autoencoder model is able to receive a real face image and compute its representation vector, which not only makes it possible to reconstruct the input image faithfully, but also separates the concepts related to specific scales. Second, we introduce a new scheme to enable generative adversarial networks to learn a representation for face images which is composed of the representations for smaller facial components. This is inspired by the fact that all face images display the same underlying structure. As a result, a face image can be divided into parts with fixed positions each containing specific facial components only. Learning a separate distribution for each of these parts is equivalent to disentangling these components in the representation space.","abstract_html":"Among the different categories of natural images, face images are very important because of the role they play in human social interactions. It is recognised that despite all the recent advances of artificial intelligence using deep neural networks, computers are still struggling at achieving a rich and flexible understanding of face images comparable to humans&amp;apos; face perception abilities. This thesis aims at finding fully unsupervised ways for learning a transformation from face images pixel space to a representation space in which the underlying facial concepts are captured and disentangled. We propose that it is possible to utilize clues from the real 3D world in order to guide the representation learner in the direction of disentangling facial concepts. We conduct two studies in order to test this hypothesis. First, we propose a deep autoencoder model for extracting facial concepts based on their scales. We introduce an adaptive resolution reconstruction loss inspired by the fact that different categories of concepts are encoded in (and can be captured from) different resolutions of face images. With the help of this new reconstruction loss, the deep autoencoder model is able to receive a real face image and compute its representation vector, which not only makes it possible to reconstruct the input image faithfully, but also separates the concepts related to specific scales. Second, we introduce a new scheme to enable generative adversarial networks to learn a representation for face images which is composed of the representations for smaller facial components. This is inspired by the fact that all face images display the same underlying structure. As a result, a face image can be divided into parts with fixed positions each containing specific facial components only. Learning a separate distribution for each of these parts is equivalent to disentangling these components in the representation space.","abstract_has_math":false,"creators":["Abdolahnejad Bahramabadi, Mahla"],"institution":"Carleton University","degree_name":"Doctor of Philosophy (Ph.D.)","degree_level":"Doctoral","degree_discipline":"Engineering, Electrical and Computer","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022","date_published":"2022","updated_at":"2026-07-24T01:34:38Z","subjects":[],"languages":["en"],"rights":["Copyright © 2022 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, research, scholarship, and teaching. Theses may only be shared by linking to Carleton University Institutional Repository and no part may be used without proper attribution to the author. No part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.22215/etd/2022-15326"],"render_values":[{"text":"10.22215/etd/2022-15326","href":"https://doi.org/10.22215/etd/2022-15326","code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/20.500.14718/42801","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Abdolahnejad Bahramabadi, Mahla"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-04-08T20:44:18Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-04-08T20:44:18Z"]},{"key":"dc:date.issued","label":"Date","values":["2022"]},{"key":"dc:publisher","label":"Institution","values":["Carleton University"]},{"key":"dc:type","label":"Dc Type","values":["thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Engineering, Electrical and Computer"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (Ph.D.)"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright © 2022 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, research, scholarship, and teaching. Theses may only be shared by linking to Carleton University Institutional Repository and no part may be used without proper attribution to the author. No part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.22215/etd/2022-15326"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/20.500.14718/42801"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Among the different categories of natural images, face images are very important because of the role they play in human social interactions. It is recognised that despite all the recent advances of artificial intelligence using deep neural networks, computers are still struggling at achieving a rich and flexible understanding of face images comparable to humans&apos; face perception abilities. This thesis aims at finding fully unsupervised ways for learning a transformation from face images pixel space to a representation space in which the underlying facial concepts are captured and disentangled. We propose that it is possible to utilize clues from the real 3D world in order to guide the representation learner in the direction of disentangling facial concepts. We conduct two studies in order to test this hypothesis. First, we propose a deep autoencoder model for extracting facial concepts based on their scales. We introduce an adaptive resolution reconstruction loss inspired by the fact that different categories of concepts are encoded in (and can be captured from) different resolutions of face images. With the help of this new reconstruction loss, the deep autoencoder model is able to receive a real face image and compute its representation vector, which not only makes it possible to reconstruct the input image faithfully, but also separates the concepts related to specific scales. Second, we introduce a new scheme to enable generative adversarial networks to learn a representation for face images which is composed of the representations for smaller facial components. This is inspired by the fact that all face images display the same underlying structure. As a result, a face image can be divided into parts with fixed positions each containing specific facial components only. Learning a separate distribution for each of these parts is equivalent to disentangling these components in the representation space."]},{"key":"dc:title","label":"Title","values":["Deep Generative Models for Unsupervised Scale-Based and Position-Based Disentanglement of Concepts from Face Images."]}]}],"canonical_facts":{"dc:creator":["Abdolahnejad Bahramabadi, Mahla"],"dc:date.accessioned":["2025-04-08T20:44:18Z"],"dc:date.available":["2025-04-08T20:44:18Z"],"dc:date.issued":["2022"],"dc:description.abstract":["Among the different categories of natural images, face images are very important because of the role they play in human social interactions. It is recognised that despite all the recent advances of artificial intelligence using deep neural networks, computers are still struggling at achieving a rich and flexible understanding of face images comparable to humans&apos; face perception abilities. This thesis aims at finding fully unsupervised ways for learning a transformation from face images pixel space to a representation space in which the underlying facial concepts are captured and disentangled. We propose that it is possible to utilize clues from the real 3D world in order to guide the representation learner in the direction of disentangling facial concepts. We conduct two studies in order to test this hypothesis. First, we propose a deep autoencoder model for extracting facial concepts based on their scales. We introduce an adaptive resolution reconstruction loss inspired by the fact that different categories of concepts are encoded in (and can be captured from) different resolutions of face images. With the help of this new reconstruction loss, the deep autoencoder model is able to receive a real face image and compute its representation vector, which not only makes it possible to reconstruct the input image faithfully, but also separates the concepts related to specific scales. Second, we introduce a new scheme to enable generative adversarial networks to learn a representation for face images which is composed of the representations for smaller facial components. This is inspired by the fact that all face images display the same underlying structure. As a result, a face image can be divided into parts with fixed positions each containing specific facial components only. Learning a separate distribution for each of these parts is equivalent to disentangling these components in the representation space."],"dc:identifier.doi":["10.22215/etd/2022-15326"],"dc:identifier.uri":["https://hdl.handle.net/20.500.14718/42801"],"dc:language.iso":["en"],"dc:publisher":["Carleton University"],"dc:rights":["Copyright © 2022 the author(s). Theses may be used for non-commercial research, educational, or related academic purposes only. Such uses include personal study, research, scholarship, and teaching. Theses may only be shared by linking to Carleton University Institutional Repository and no part may be used without proper attribution to the author. No part may be used for commercial purposes directly or indirectly via a for-profit platform; no adaptation or derivative works are permitted without consent from the copyright owner."],"dc:title":["Deep Generative Models for Unsupervised Scale-Based and Position-Based Disentanglement of Concepts from Face Images."],"dc:type":["thesis"],"thesis:degree_discipline":["Engineering, Electrical and Computer"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy (Ph.D.)"]},"updated_at":"2026-07-24T01:34:38Z"}