{"id":{"repo_id":"potsdam-thes","oai_identifier":"oai:kobv.de-opus4-uni-potsdam:67962"},"canonical_url":"https://search.dev.ndltd.org/etd/potsdam-thes/oai:kobv.de-opus4-uni-potsdam:67962","repository":{"repo_id":"potsdam-thes","name":"Universität Potsdam - Thes","base_url":"https://publishup.uni-potsdam.de/opus4-ubp/oai"},"display":{"title":"Study of the Diffusion Map method in the context of social science data sets","abstract":"The Diffusion Map is a nonlinear dimensionality reduction technique used to analyze high-dimensional data, with recent applications extending to datasets from the social sciences. Previous research has given little attention to how the specific characteristics of these datasets might influence the results of the Diffusion Map and what conditions must be met for the Diffusion Map to yield meaningful and interpretable results. Moreover, there is a lack of clear, comprehensive explanations of the fundamental principles, which has led to misunderstandings in the literature. This work first addresses the fundamental principles of the Diffusion Map and compares them with other spectral methods. It investigates the impact of the Diffusion Map parameters as well as the structure of the underlying data on the results. The V-Dem democracy dataset, British census data, and data on German urban and rural districts are then analyzed, considering their possible natural parameters. A focus is placed on the benefits of the Diffusion Map in comparison to the established linear principal component analysis (PCA). The analysis shows that the time parameter t of the Diffusion Map framework has no significant influence on the analysis. In contrast, discrete and redundant variables, as well as the scaling and normalization of the data, have a substantial impact. Unlike PCA, the Diffusion Map eigenspectrum does not provide a clear indication of which components are important. Therefore, typical polynomial patterns related to one-dimensional datasets within the Diffusion Map framework are explored. The thesis presents insights suggesting that several underconsidered effects need further examination, and emphasizes the need for a framework to accurately analyze complex datasets using the Diffusion Map.","abstract_html":"The Diffusion Map is a nonlinear dimensionality reduction technique used to analyze high-dimensional data, with recent applications extending to datasets from the social sciences. Previous research has given little attention to how the specific characteristics of these datasets might influence the results of the Diffusion Map and what conditions must be met for the Diffusion Map to yield meaningful and interpretable results. Moreover, there is a lack of clear, comprehensive explanations of the fundamental principles, which has led to misunderstandings in the literature. This work first addresses the fundamental principles of the Diffusion Map and compares them with other spectral methods. It investigates the impact of the Diffusion Map parameters as well as the structure of the underlying data on the results. The V-Dem democracy dataset, British census data, and data on German urban and rural districts are then analyzed, considering their possible natural parameters. A focus is placed on the benefits of the Diffusion Map in comparison to the established linear principal component analysis (PCA). The analysis shows that the time parameter t of the Diffusion Map framework has no significant influence on the analysis. In contrast, discrete and redundant variables, as well as the scaling and normalization of the data, have a substantial impact. Unlike PCA, the Diffusion Map eigenspectrum does not provide a clear indication of which components are important. Therefore, typical polynomial patterns related to one-dimensional datasets within the Diffusion Map framework are explored. The thesis presents insights suggesting that several underconsidered effects need further examination, and emphasizes the need for a framework to accurately analyze complex datasets using the Diffusion Map.","abstract_has_math":false,"creators":["Beier, Sönke"],"institution":"Universität Potsdam","degree_name":null,"degree_level":"master","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Pirker-Díaz, Paula","Wiesner, Karoline","Olbrich, Eckehard"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-12","date_published":"2025-03-12","updated_at":"2026-07-24T03:52:13Z","subjects":["Diffusion Map","dimensionality reduction","manifold learning","census data","varieties of democracy (V-Dem)","spectral embedding","sociophysics","complexity science","application Diffusion Map","Zensus","Komplexitätswissenschaften","Dimensionsreduktion","Soziophysik","Demokratieindex"],"languages":[],"rights":["CC-BY - Namensnennung 4.0 International"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://publishup.uni-potsdam.de/frontdoor/index/index/docId/67962","outbound_label":"Repository record","outbound_source":"source_url"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Pirker-Díaz, Paula","Wiesner, Karoline","Olbrich, Eckehard"]},{"key":"dc:creator","label":"Author","values":["Beier, Sönke"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:publisher","label":"Institution","values":["Universität Potsdam"]},{"key":"dc:type","label":"Dc Type","values":["masterThesis"]},{"key":"thesis:degree_level","label":"Degree Level","values":["master"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Universität Potsdam"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Diffusion Map","dimensionality reduction","manifold learning","census data","varieties of democracy (V-Dem)","spectral embedding","sociophysics","complexity science","application Diffusion Map","Zensus","Komplexitätswissenschaften","Dimensionsreduktion","Soziophysik","Demokratieindex"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["CC-BY - Namensnennung 4.0 International"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The Diffusion Map is a nonlinear dimensionality reduction technique used to analyze high-dimensional data, with recent applications extending to datasets from the social sciences. Previous research has given little attention to how the specific characteristics of these datasets might influence the results of the Diffusion Map and what conditions must be met for the Diffusion Map to yield meaningful and interpretable results. Moreover, there is a lack of clear, comprehensive explanations of the fundamental principles, which has led to misunderstandings in the literature. This work first addresses the fundamental principles of the Diffusion Map and compares them with other spectral methods. It investigates the impact of the Diffusion Map parameters as well as the structure of the underlying data on the results. The V-Dem democracy dataset, British census data, and data on German urban and rural districts are then analyzed, considering their possible natural parameters. A focus is placed on the benefits of the Diffusion Map in comparison to the established linear principal component analysis (PCA). The analysis shows that the time parameter t of the Diffusion Map framework has no significant influence on the analysis. In contrast, discrete and redundant variables, as well as the scaling and normalization of the data, have a substantial impact. Unlike PCA, the Diffusion Map eigenspectrum does not provide a clear indication of which components are important. Therefore, typical polynomial patterns related to one-dimensional datasets within the Diffusion Map framework are explored. The thesis presents insights suggesting that several underconsidered effects need further examination, and emphasizes the need for a framework to accurately analyze complex datasets using the Diffusion Map.","In der heutigen Zeit werden zunehmend hochdimensionale Datensätze erzeugt und analysiert. Die Diffusion Map ist eine nicht-lineare Dimensionsreduktionsmethode, die dazu dient, solche Datensätze greifbar zu machen. In den letzten Jahren fand diese Methode auch Anwendung bei der Analyse von Datensätzen aus den Sozialwissenschaften. Bisherige Arbeiten haben jedoch wenig untersucht, welchen Einfluss die besonderen Eigenschaften dieser Datensätze auf die Ergebnisse der Diffusion Map haben können und welche Voraussetzungen erfüllt sein müssen, damit die Diffusion Map sinnvolle und interpretierbare Ergebnisse liefert. Zudem mangelt es an klaren, zusammenfassenden Erklärungen der Grundprinzipien, was zu Missverständnissen in der Literatur führt. Diese Arbeit setzt sich zunächst mit den Grundprinzipien der Diffusion Map auseinander und vergleicht diese mit anderen spektralen Methoden. Es wird untersucht, welche Einflüsse die Parameter der Diffusion Map sowie die Struktur der zugrunde liegenden Daten auf die Ergebnisse haben. Anschließend werden unter Berücksichtigung dieses Wissens der V-Dem Demokratiedatensatz, britische Census-Daten und Daten über deutsche Städte und Landkreise analysiert. Ein besonderer Fokus liegt auf dem Nutzen der Diffusion Map im Vergleich zur etablierten linearen Principal Component Analysis (PCA). Die Analyse zeigt, dass der Zeitparameter t keinen nennenswerten Einfluss auf das Ergebnis hat. Hingegen haben diskrete und redundante Variablen sowie Skalierungen und Normalisierungen der Daten einen erheblichen Einfluss. Im Gegensatz zur PCA bieten die Diffusion Map Eigenwerte keine klare Auskunft darüber, welche Dimensionen wichtig sind. Es werden die möglichen natürlichen Parameter der drei untersuchten Datensätze untersucht, typische Formen der Diffusion Map identifiziert und gezeigt, dass auch die PCA die Daten sinnvoll anordnen kann. Die Arbeit kommt zu dem Schluss, dass verschiedene Effekte der Diffusion Map noch nicht ausreichend erforscht sind. Es wird ein Framework angeregt, das es ermöglicht, Datensätze mit der Diffusion Map korrekt zu analysieren und Interpretationsfehler zu vermeiden. Diese Arbeit liefert erste Impulse für die Entwicklung eines solchen Ansatzes."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Study of the Diffusion Map method in the context of social science data sets","Untersuchung der 'Diffusion Map' Methode im Kontext sozialwissenschaftlicher Datensätze"]}]}],"canonical_facts":{"dc:contributor":["Pirker-Díaz, Paula","Wiesner, Karoline","Olbrich, Eckehard"],"dc:creator":["Beier, Sönke"],"dc:description.abstract":["The Diffusion Map is a nonlinear dimensionality reduction technique used to analyze high-dimensional data, with recent applications extending to datasets from the social sciences. Previous research has given little attention to how the specific characteristics of these datasets might influence the results of the Diffusion Map and what conditions must be met for the Diffusion Map to yield meaningful and interpretable results. Moreover, there is a lack of clear, comprehensive explanations of the fundamental principles, which has led to misunderstandings in the literature. This work first addresses the fundamental principles of the Diffusion Map and compares them with other spectral methods. It investigates the impact of the Diffusion Map parameters as well as the structure of the underlying data on the results. The V-Dem democracy dataset, British census data, and data on German urban and rural districts are then analyzed, considering their possible natural parameters. A focus is placed on the benefits of the Diffusion Map in comparison to the established linear principal component analysis (PCA). The analysis shows that the time parameter t of the Diffusion Map framework has no significant influence on the analysis. In contrast, discrete and redundant variables, as well as the scaling and normalization of the data, have a substantial impact. Unlike PCA, the Diffusion Map eigenspectrum does not provide a clear indication of which components are important. Therefore, typical polynomial patterns related to one-dimensional datasets within the Diffusion Map framework are explored. The thesis presents insights suggesting that several underconsidered effects need further examination, and emphasizes the need for a framework to accurately analyze complex datasets using the Diffusion Map.","In der heutigen Zeit werden zunehmend hochdimensionale Datensätze erzeugt und analysiert. Die Diffusion Map ist eine nicht-lineare Dimensionsreduktionsmethode, die dazu dient, solche Datensätze greifbar zu machen. In den letzten Jahren fand diese Methode auch Anwendung bei der Analyse von Datensätzen aus den Sozialwissenschaften. Bisherige Arbeiten haben jedoch wenig untersucht, welchen Einfluss die besonderen Eigenschaften dieser Datensätze auf die Ergebnisse der Diffusion Map haben können und welche Voraussetzungen erfüllt sein müssen, damit die Diffusion Map sinnvolle und interpretierbare Ergebnisse liefert. Zudem mangelt es an klaren, zusammenfassenden Erklärungen der Grundprinzipien, was zu Missverständnissen in der Literatur führt. Diese Arbeit setzt sich zunächst mit den Grundprinzipien der Diffusion Map auseinander und vergleicht diese mit anderen spektralen Methoden. Es wird untersucht, welche Einflüsse die Parameter der Diffusion Map sowie die Struktur der zugrunde liegenden Daten auf die Ergebnisse haben. Anschließend werden unter Berücksichtigung dieses Wissens der V-Dem Demokratiedatensatz, britische Census-Daten und Daten über deutsche Städte und Landkreise analysiert. Ein besonderer Fokus liegt auf dem Nutzen der Diffusion Map im Vergleich zur etablierten linearen Principal Component Analysis (PCA). Die Analyse zeigt, dass der Zeitparameter t keinen nennenswerten Einfluss auf das Ergebnis hat. Hingegen haben diskrete und redundante Variablen sowie Skalierungen und Normalisierungen der Daten einen erheblichen Einfluss. Im Gegensatz zur PCA bieten die Diffusion Map Eigenwerte keine klare Auskunft darüber, welche Dimensionen wichtig sind. Es werden die möglichen natürlichen Parameter der drei untersuchten Datensätze untersucht, typische Formen der Diffusion Map identifiziert und gezeigt, dass auch die PCA die Daten sinnvoll anordnen kann. Die Arbeit kommt zu dem Schluss, dass verschiedene Effekte der Diffusion Map noch nicht ausreichend erforscht sind. Es wird ein Framework angeregt, das es ermöglicht, Datensätze mit der Diffusion Map korrekt zu analysieren und Interpretationsfehler zu vermeiden. Diese Arbeit liefert erste Impulse für die Entwicklung eines solchen Ansatzes."],"dc:format.medium":["application/pdf"],"dc:publisher":["Universität Potsdam"],"dc:rights":["CC-BY - Namensnennung 4.0 International"],"dc:subject":["Diffusion Map","dimensionality reduction","manifold learning","census data","varieties of democracy (V-Dem)","spectral embedding","sociophysics","complexity science","application Diffusion Map","Zensus","Komplexitätswissenschaften","Dimensionsreduktion","Soziophysik","Demokratieindex"],"dc:title":["Study of the Diffusion Map method in the context of social science data sets","Untersuchung der 'Diffusion Map' Methode im Kontext sozialwissenschaftlicher Datensätze"],"dc:type":["masterThesis"],"thesis:degree_level":["master"],"thesis:institution_name":["Universität Potsdam"]},"updated_at":"2026-07-24T03:52:13Z"}