Back to results

Universidad de Cadiz

Navigating Diverse Datasets in the Face of Uncertainty

Abstract

dc:description.abstract

When exploring big volumes of data, one of the challenging aspects is their diversity of origin. Multiple files that have not yet been ingested into a database system may contain information of interest to a researcher, who must curate, understand and sieve their content before being able to extract knowledge. Performance is one of the greatest difficulties in exploring these datasets. On the one hand, examining non-indexed, unprocessed files can be inefficient. On the other hand, any processing before its understanding introduces latency and potentially un- necessary work if the chosen schema matches poorly the data. We have surveyed the state-of-the-art and, fortunately, there exist multiple proposal of solutions to handle data in-situ performantly. Another major difficulty is matching files from multiple origins since their schema and layout may not be compatible or properly documented. Most surveyed solutions overlook this problem, especially for numeric, uncertain data, as is typical in fields like astronomy. The main objective of our research is to assist data scientists during the exploration of unprocessed, numerical, raw data distributed across multiple files based solely on its intrinsic distribution. In this thesis, we first introduce the concept of Equally-Distributed Dependencies, which provides the foundations to match this kind of dataset. We propose PresQ, a novel algorithm that finds quasi-cliques on hypergraphs based on their expected statistical properties. The probabilistic approach of PresQ can be successfully exploited to mine EDD between diverse datasets when the underlying populations can be assumed to be the same. Finally, we propose a two-sample statistical test based on Self-Organizing Maps (SOM). This method can outperform, in terms of power, other classifier-based two- sample tests, being in some cases comparable to kernel-based methods, with the advantage of being interpretable. Both PresQ and the SOM-based statistical test can provide insights that drive serendipitous discoveries.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Álvarez Ayllón, Alejandro
Advisors dc:contributor.advisor
  • Palomo Duarte, Manuel
  • Dodero Beardo, Juan Manuel

Subjects

dc:subject × 3

Rights

dc:rights
Statement dc:rights
  • Atribución 4.0 Internacional
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/10498/29364
OAI identifier oai:identifier
oai:rodin.uca.es:10498/29364

Chain of custody

source
Harvested from
Universidad de Cadiz
Base URL
rodin.uca.es/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Álvarez Ayllón, Alejandro. Navigating Diverse Datasets in the Face of Uncertainty. 2023. http://hdl.handle.net/10498/29364