{"id":{"repo_id":"tu-berlin","oai_identifier":"oai:depositonce.tu-berlin.de:11303/24941"},"canonical_url":"https://search.dev.ndltd.org/etd/tu-berlin/oai:depositonce.tu-berlin.de:11303/24941","repository":{"repo_id":"tu-berlin","name":"Technische Universität Berlin","base_url":"https://api-depositonce.tu-berlin.de/server/oai/request"},"display":{"title":"Knowledge-augmented and context-sensitive face perception","abstract":"Facial expressions play a crucial role in human communication, conveying a significant amount of information about an individual’s internal emotional state. In the past, psychological research has focused on facial expressions in isolation. However, recent findings in psychology and cognitive science have shifted towards highlighting the inherently contextualized nature of human perception and cognition, demonstrating that context and prior knowledge substantially influence how people perceive and behave. For instance, how someone perceives a smile changes when accompanied by affective biographical information of the respective person. Consequently, it is essential to incorporate knowledge and contextual information in a dynamic manner and inspired by human perception, into synthetic facial expression recognition (FER), to avoid a mismatch between artificial systems and human perception. Close collaboration between researchers from psychology and computer science is therefore necessary. In this thesis, we address FER using deep neural networks from an interdisciplinary, human-centered perspective with the goal of adapting it to the dynamics of human facial expression perception and enhance and enable communication between machines and humans. To facilitate communication between the different fields, we provide a joint review on existing literature in light of a novel framework and common terminology. Within this framework, we identify hallmarks of context-sensitive and knowledge-augmented human perception and deduce programmatic concepts for synthetic FER systems. We propose two distinct methods for achieving context-sensitivity: a self-supervised learning approach that clusters the latent space during pretraining using context information from additional audio and text modalities. A second method dynamically clusters facial features based on audio context during downstream inference. We evaluate both methods on state-of-the-art FER datasets and demonstrate that our approaches outperform many competitors while either being trained unsupervised without labels or exhibiting additional generative capabilities. The latter aims at realizing the ability to produce high- quality approximations of the mental representations that humans retain of the perceived expressions. Automatic generation, as we achieve it in our work, will be beneficial in research on human face perception and beyond. Additionally, simultaneous recognition and generation of expressions constitutes a tool of explainability in artificial intelligence (AI) models and allows us to compare machine and human facial representations.","abstract_html":"Facial expressions play a crucial role in human communication, conveying a significant amount of information about an individual’s internal emotional state. In the past, psychological research has focused on facial expressions in isolation. However, recent findings in psychology and cognitive science have shifted towards highlighting the inherently contextualized nature of human perception and cognition, demonstrating that context and prior knowledge substantially influence how people perceive and behave. For instance, how someone perceives a smile changes when accompanied by affective biographical information of the respective person. Consequently, it is essential to incorporate knowledge and contextual information in a dynamic manner and inspired by human perception, into synthetic facial expression recognition (FER), to avoid a mismatch between artificial systems and human perception. Close collaboration between researchers from psychology and computer science is therefore necessary. In this thesis, we address FER using deep neural networks from an interdisciplinary, human-centered perspective with the goal of adapting it to the dynamics of human facial expression perception and enhance and enable communication between machines and humans. To facilitate communication between the different fields, we provide a joint review on existing literature in light of a novel framework and common terminology. Within this framework, we identify hallmarks of context-sensitive and knowledge-augmented human perception and deduce programmatic concepts for synthetic FER systems. We propose two distinct methods for achieving context-sensitivity: a self-supervised learning approach that clusters the latent space during pretraining using context information from additional audio and text modalities. A second method dynamically clusters facial features based on audio context during downstream inference. We evaluate both methods on state-of-the-art FER datasets and demonstrate that our approaches outperform many competitors while either being trained unsupervised without labels or exhibiting additional generative capabilities. The latter aims at realizing the ability to produce high- quality approximations of the mental representations that humans retain of the perceived expressions. Automatic generation, as we achieve it in our work, will be beneficial in research on human face perception and beyond. Additionally, simultaneous recognition and generation of expressions constitutes a tool of explainability in artificial intelligence (AI) models and allows us to compare machine and human facial representations.","abstract_has_math":false,"creators":["Blume, Florian"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Hellwich, Olaf"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-27T21:28:26Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":["https://creativecommons.org/licenses/by/4.0/"],"identifier_entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://doi.org/10.14279/depositonce-23757"],"render_values":[{"text":"https://doi.org/10.14279/depositonce-23757","href":"https://doi.org/10.14279/depositonce-23757","code":true}]}]},"links":{"outbound_url":"https://depositonce.tu-berlin.de/handle/11303/24941","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Hellwich, Olaf"]},{"key":"dc:creator","label":"Author","values":["Blume, Florian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-06-11T12:56:43Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-06-11T12:56:43Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Doctoral Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights.uri","label":"Rights URI","values":["https://creativecommons.org/licenses/by/4.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://depositonce.tu-berlin.de/handle/11303/24941","https://doi.org/10.14279/depositonce-23757"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Facial expressions play a crucial role in human communication, conveying a significant amount of information about an individual’s internal emotional state. In the past, psychological research has focused on facial expressions in isolation. However, recent findings in psychology and cognitive science have shifted towards highlighting the inherently contextualized nature of human perception and cognition, demonstrating that context and prior knowledge substantially influence how people perceive and behave. For instance, how someone perceives a smile changes when accompanied by affective biographical information of the respective person. Consequently, it is essential to incorporate knowledge and contextual information in a dynamic manner and inspired by human perception, into synthetic facial expression recognition (FER), to avoid a mismatch between artificial systems and human perception. Close collaboration between researchers from psychology and computer science is therefore necessary. In this thesis, we address FER using deep neural networks from an interdisciplinary, human-centered perspective with the goal of adapting it to the dynamics of human facial expression perception and enhance and enable communication between machines and humans. To facilitate communication between the different fields, we provide a joint review on existing literature in light of a novel framework and common terminology. Within this framework, we identify hallmarks of context-sensitive and knowledge-augmented human perception and deduce programmatic concepts for synthetic FER systems. We propose two distinct methods for achieving context-sensitivity: a self-supervised learning approach that clusters the latent space during pretraining using context information from additional audio and text modalities. A second method dynamically clusters facial features based on audio context during downstream inference. We evaluate both methods on state-of-the-art FER datasets and demonstrate that our approaches outperform many competitors while either being trained unsupervised without labels or exhibiting additional generative capabilities. The latter aims at realizing the ability to produce high- quality approximations of the mental representations that humans retain of the perceived expressions. Automatic generation, as we achieve it in our work, will be beneficial in research on human face perception and beyond. Additionally, simultaneous recognition and generation of expressions constitutes a tool of explainability in artificial intelligence (AI) models and allows us to compare machine and human facial representations.","Gesichtsausdrücke spielen in der menschlichen Kommunikation eine entscheidende Rolle, denn sie vermitteln essenzielle Informationen über den inneren emotionalen Zustand einer Person. In der Vergangenheit hat die psychologische Forschung Gesichtsausdrücke primär in Isolation betrachtet. In der neueren Forschung gab es jedoch einen Paradigmenwechsel hin zu der Erkenntnis, dass die menschliche Wahrnehmung und menschliche Denkprozesse stark von Kontext und (Vor-)Wissen beeinflusst sind. So ändert sich zum Beispiel mit biografischen Informationen über die Person, wie deren Lächeln wahrgenommen wird. Daher ist es unerlässlich, Wissen und Kontextinformationen, dynamisch und inspiriert durch menschliche Wahrnehmung, in Gesichtsausdruckserkennung (GAE) einzuarbeiten, um eine Diskrepanz zwischen künstlicher Verarbeitung und menschlicher Wahrnehmung zu verhindern. Eine enge Zusammenarbeit von Forschern in Psychologie und Informatik ist notwendig. In dieser Dissertation arbeiten wir daran, Gesichtsausdrücke automatisch mit tiefen neuronalen Netzen zu erkennen, aus einer interdisziplinären und auf den Menschen fokussierten Perspektive heraus mit dem Ziel, GAE an die Dynamiken menschlicher Gesichtsausdruckswahrnehmung zu adaptieren und somit die Kommunikation zwischen Mensch und Maschine zu erleichtern und zu ermöglichen. Um die Kommunikation zwischen allen Feldern zu ermöglichen, begutachten wir die existierende Literatur unter Bezugnahme auf ein von uns entwickeltes Framework und einer gemeinsamen Terminologie. In diesem Framework identifizieren wir Kennzeichen menschlicher, kontextualisierter und wissenserweiterter Wahrnehmung und deduzieren programmatische Konzepte für synthetische GAE Systeme. Wir stellen zwei verschiedene Methoden vor, um Kontextsensitivität in GAE zu realisieren: Ein selbst-überwachter Lernansatz, der den Repräsentationsraum unter Zuhilfenahme von Kontextinformation, in Form von zusätzlichen Audio- und Textmodalitäten strukturiert. Eine zweite Methode adaptiert dynamisch die Repräsentation des Gesichtsausdrucks anhand von Audiokontext während der Inferenz. Wir evaluieren beide Methoden auf aktuellen GAE Datensätzen und demonstrieren, dass unsere Ansätze Konkurrenzmethoden übertreffen, wobei unsere entweder unüberwacht ohne Klassen trainiert werden, oder aber mit zusätzlichen generativen Fähigkeiten ausgestattet sind. Letzteres erlaubt Annäherungen an mentale Abbildungen von Gesichtsausdrücken zu generieren. Automatische Generierung, wie wir sie in der vorliegenden Arbeit erreichen, kann für die Forschung mit Gesichtern nützlich sein. Gleichzeitig erhalten wir so ein Werkzeug für Erklärbarkeit für künstliche Intelligenz (KI) Systeme und können die synthetischen Repräsentationen von Gesichtern mit menschlichen vergleichen."]},{"key":"dc:title","label":"Title","values":["Knowledge-augmented and context-sensitive face perception"]}]}],"canonical_facts":{"dc:contributor.advisor":["Hellwich, Olaf"],"dc:creator":["Blume, Florian"],"dc:date.accessioned":["2025-06-11T12:56:43Z"],"dc:date.available":["2025-06-11T12:56:43Z"],"dc:date.issued":["2025"],"dc:description.abstract":["Facial expressions play a crucial role in human communication, conveying a significant amount of information about an individual’s internal emotional state. In the past, psychological research has focused on facial expressions in isolation. However, recent findings in psychology and cognitive science have shifted towards highlighting the inherently contextualized nature of human perception and cognition, demonstrating that context and prior knowledge substantially influence how people perceive and behave. For instance, how someone perceives a smile changes when accompanied by affective biographical information of the respective person. Consequently, it is essential to incorporate knowledge and contextual information in a dynamic manner and inspired by human perception, into synthetic facial expression recognition (FER), to avoid a mismatch between artificial systems and human perception. Close collaboration between researchers from psychology and computer science is therefore necessary. In this thesis, we address FER using deep neural networks from an interdisciplinary, human-centered perspective with the goal of adapting it to the dynamics of human facial expression perception and enhance and enable communication between machines and humans. To facilitate communication between the different fields, we provide a joint review on existing literature in light of a novel framework and common terminology. Within this framework, we identify hallmarks of context-sensitive and knowledge-augmented human perception and deduce programmatic concepts for synthetic FER systems. We propose two distinct methods for achieving context-sensitivity: a self-supervised learning approach that clusters the latent space during pretraining using context information from additional audio and text modalities. A second method dynamically clusters facial features based on audio context during downstream inference. We evaluate both methods on state-of-the-art FER datasets and demonstrate that our approaches outperform many competitors while either being trained unsupervised without labels or exhibiting additional generative capabilities. The latter aims at realizing the ability to produce high- quality approximations of the mental representations that humans retain of the perceived expressions. Automatic generation, as we achieve it in our work, will be beneficial in research on human face perception and beyond. Additionally, simultaneous recognition and generation of expressions constitutes a tool of explainability in artificial intelligence (AI) models and allows us to compare machine and human facial representations.","Gesichtsausdrücke spielen in der menschlichen Kommunikation eine entscheidende Rolle, denn sie vermitteln essenzielle Informationen über den inneren emotionalen Zustand einer Person. In der Vergangenheit hat die psychologische Forschung Gesichtsausdrücke primär in Isolation betrachtet. In der neueren Forschung gab es jedoch einen Paradigmenwechsel hin zu der Erkenntnis, dass die menschliche Wahrnehmung und menschliche Denkprozesse stark von Kontext und (Vor-)Wissen beeinflusst sind. So ändert sich zum Beispiel mit biografischen Informationen über die Person, wie deren Lächeln wahrgenommen wird. Daher ist es unerlässlich, Wissen und Kontextinformationen, dynamisch und inspiriert durch menschliche Wahrnehmung, in Gesichtsausdruckserkennung (GAE) einzuarbeiten, um eine Diskrepanz zwischen künstlicher Verarbeitung und menschlicher Wahrnehmung zu verhindern. Eine enge Zusammenarbeit von Forschern in Psychologie und Informatik ist notwendig. In dieser Dissertation arbeiten wir daran, Gesichtsausdrücke automatisch mit tiefen neuronalen Netzen zu erkennen, aus einer interdisziplinären und auf den Menschen fokussierten Perspektive heraus mit dem Ziel, GAE an die Dynamiken menschlicher Gesichtsausdruckswahrnehmung zu adaptieren und somit die Kommunikation zwischen Mensch und Maschine zu erleichtern und zu ermöglichen. Um die Kommunikation zwischen allen Feldern zu ermöglichen, begutachten wir die existierende Literatur unter Bezugnahme auf ein von uns entwickeltes Framework und einer gemeinsamen Terminologie. In diesem Framework identifizieren wir Kennzeichen menschlicher, kontextualisierter und wissenserweiterter Wahrnehmung und deduzieren programmatische Konzepte für synthetische GAE Systeme. Wir stellen zwei verschiedene Methoden vor, um Kontextsensitivität in GAE zu realisieren: Ein selbst-überwachter Lernansatz, der den Repräsentationsraum unter Zuhilfenahme von Kontextinformation, in Form von zusätzlichen Audio- und Textmodalitäten strukturiert. Eine zweite Methode adaptiert dynamisch die Repräsentation des Gesichtsausdrucks anhand von Audiokontext während der Inferenz. Wir evaluieren beide Methoden auf aktuellen GAE Datensätzen und demonstrieren, dass unsere Ansätze Konkurrenzmethoden übertreffen, wobei unsere entweder unüberwacht ohne Klassen trainiert werden, oder aber mit zusätzlichen generativen Fähigkeiten ausgestattet sind. Letzteres erlaubt Annäherungen an mentale Abbildungen von Gesichtsausdrücken zu generieren. Automatische Generierung, wie wir sie in der vorliegenden Arbeit erreichen, kann für die Forschung mit Gesichtern nützlich sein. Gleichzeitig erhalten wir so ein Werkzeug für Erklärbarkeit für künstliche Intelligenz (KI) Systeme und können die synthetischen Repräsentationen von Gesichtern mit menschlichen vergleichen."],"dc:identifier.uri":["https://depositonce.tu-berlin.de/handle/11303/24941","https://doi.org/10.14279/depositonce-23757"],"dc:language.iso":["en"],"dc:rights.uri":["https://creativecommons.org/licenses/by/4.0/"],"dc:title":["Knowledge-augmented and context-sensitive face perception"],"dc:type":["Doctoral Thesis"]},"updated_at":"2026-07-27T21:28:26Z"}