Technische Universität Berlin
Knowledge-augmented and context-sensitive face perception
Abstract
dc:description.abstractFacial expressions play a crucial role in human communication, conveying a significant amount of information about an individual’s internal emotional state. In the past, psychological research has focused on facial expressions in isolation. However, recent findings in psychology and cognitive science have shifted towards highlighting the inherently contextualized nature of human perception and cognition, demonstrating that context and prior knowledge substantially influence how people perceive and behave. For instance, how someone perceives a smile changes when accompanied by affective biographical information of the respective person. Consequently, it is essential to incorporate knowledge and contextual information in a dynamic manner and inspired by human perception, into synthetic facial expression recognition (FER), to avoid a mismatch between artificial systems and human perception. Close collaboration between researchers from psychology and computer science is therefore necessary. In this thesis, we address FER using deep neural networks from an interdisciplinary, human-centered perspective with the goal of adapting it to the dynamics of human facial expression perception and enhance and enable communication between machines and humans. To facilitate communication between the different fields, we provide a joint review on existing literature in light of a novel framework and common terminology. Within this framework, we identify hallmarks of context-sensitive and knowledge-augmented human perception and deduce programmatic concepts for synthetic FER systems. We propose two distinct methods for achieving context-sensitivity: a self-supervised learning approach that clusters the latent space during pretraining using context information from additional audio and text modalities. A second method dynamically clusters facial features based on audio context during downstream inference. We evaluate both methods on state-of-the-art FER datasets and demonstrate that our approaches outperform many competitors while either being trained unsupervised without labels or exhibiting additional generative capabilities. The latter aims at realizing the ability to produce high- quality approximations of the mental representations that humans retain of the perceived expressions. Automatic generation, as we achieve it in our work, will be beneficial in research on human face perception and beyond. Additionally, simultaneous recognition and generation of expressions constitutes a tool of explainability in artificial intelligence (AI) models and allows us to compare machine and human facial representations.
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Blume, Florian
- Advisor dc:contributor.advisor
-
- Hellwich, Olaf
Rights
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Identifier URI
- https://doi.org/10.14279/depositonce-23757
- OAI identifier oai:identifier
- oai:depositonce.tu-berlin.de:11303/24941