{"id":{"repo_id":"liverpool-jm","oai_identifier":"oai:researchonline.ljmu.ac.uk:10869"},"canonical_url":"https://search.dev.ndltd.org/etd/liverpool-jm/oai:researchonline.ljmu.ac.uk:10869","repository":{"repo_id":"liverpool-jm","name":"Liverpool Jon Moores University","base_url":"https://researchonline.ljmu.ac.uk/cgi/oai2"},"display":{"title":"Identification of Data Structure with Machine Learning: From Fisher to Bayesian networks","abstract":"This thesis proposes a theoretical framework to thoroughly analyse the structure of a dataset in terms of a) metric, b) density and c) feature associations. To look into the first aspect, Fisher's metric learning algorithms are the foundations of a novel manifold based on the information and complexity of a classification model. When looking at the density aspect, the Probabilistic Quantum clustering, a Bayesian version of the original Quantum Clustering is proposed. The clustering results will depend on local density variations, which is a desired feature when dealing with heteroscedastic data. To address the third aspect, the constraint-based PC-algorithm is the starting point of many structure learning algorithms, it is focused on finding feature associations by means of conditional independent tests. This is then used to select Bayesian networks, based on a regularized likelihood score. These three topics of data structure analysis were fully tested with synthetic data examples and real cases, which allowed us to unravel and discuss the advantages and limitations of these algorithms. One of the biggest challenges encountered was related to the application of these methods to a Big Data dataset that was analysed within the framework of a collaboration with a large UK retailer, where the interest was in the identification of the data structure underlying customer shopping baskets.","abstract_html":"This thesis proposes a theoretical framework to thoroughly analyse the structure of a dataset in terms of a) metric, b) density and c) feature associations. To look into the first aspect, Fisher&#x27;s metric learning algorithms are the foundations of a novel manifold based on the information and complexity of a classification model. When looking at the density aspect, the Probabilistic Quantum clustering, a Bayesian version of the original Quantum Clustering is proposed. The clustering results will depend on local density variations, which is a desired feature when dealing with heteroscedastic data. To address the third aspect, the constraint-based PC-algorithm is the starting point of many structure learning algorithms, it is focused on finding feature associations by means of conditional independent tests. This is then used to select Bayesian networks, based on a regularized likelihood score. These three topics of data structure analysis were fully tested with synthetic data examples and real cases, which allowed us to unravel and discuss the advantages and limitations of these algorithms. One of the biggest challenges encountered was related to the application of these methods to a Big Data dataset that was analysed within the framework of a collaboration with a large UK retailer, where the interest was in the identification of the data structure underlying customer shopping baskets.","abstract_has_math":false,"creators":["Casana Eslava, R"],"institution":"Liverpool John Moores University","degree_name":"phd","degree_level":"doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Jarman, IH","Lisboa, PJ","Ortega-Martorell, S"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2019,"date_issued":"2019-06","date_published":"2019-06","updated_at":"2026-07-24T06:30:28Z","subjects":["QA75 Electronic computers. Computer science"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.24377/LJMU.t.00010869","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jarman, IH","Lisboa, PJ","Ortega-Martorell, S"]},{"key":"dc:creator","label":"Author","values":["Casana Eslava, R"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2019-06-11"]},{"key":"dc:date.issued","label":"Date","values":["2019-06"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Computer Science and Mathematics"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["Liverpool John Moores University"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://researchonline.ljmu.ac.uk/id/eprint/10869/"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["phd"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["QA75 Electronic computers. Computer science"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.24377/LJMU.t.00010869"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://researchonline.ljmu.ac.uk/id/eprint/10869/1/2018raulcasanaphd.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis proposes a theoretical framework to thoroughly analyse the structure of a dataset in terms of a) metric, b) density and c) feature associations. To look into the first aspect, Fisher's metric learning algorithms are the foundations of a novel manifold based on the information and complexity of a classification model. When looking at the density aspect, the Probabilistic Quantum clustering, a Bayesian version of the original Quantum Clustering is proposed. The clustering results will depend on local density variations, which is a desired feature when dealing with heteroscedastic data. To address the third aspect, the constraint-based PC-algorithm is the starting point of many structure learning algorithms, it is focused on finding feature associations by means of conditional independent tests. This is then used to select Bayesian networks, based on a regularized likelihood score. These three topics of data structure analysis were fully tested with synthetic data examples and real cases, which allowed us to unravel and discuss the advantages and limitations of these algorithms. One of the biggest challenges encountered was related to the application of these methods to a Big Data dataset that was analysed within the framework of a collaboration with a large UK retailer, where the interest was in the identification of the data structure underlying customer shopping baskets."]},{"key":"dc:format","label":"Dc Format","values":["text"]},{"key":"dc:title","label":"Title","values":["Identification of Data Structure with Machine Learning: From Fisher to Bayesian networks"]}]}],"canonical_facts":{"dc:contributor":["Jarman, IH","Lisboa, PJ","Ortega-Martorell, S"],"dc:creator":["Casana Eslava, R"],"dc:date":["2019-06-11"],"dc:date.issued":["2019-06"],"dc:description.abstract":["This thesis proposes a theoretical framework to thoroughly analyse the structure of a dataset in terms of a) metric, b) density and c) feature associations. To look into the first aspect, Fisher's metric learning algorithms are the foundations of a novel manifold based on the information and complexity of a classification model. When looking at the density aspect, the Probabilistic Quantum clustering, a Bayesian version of the original Quantum Clustering is proposed. The clustering results will depend on local density variations, which is a desired feature when dealing with heteroscedastic data. To address the third aspect, the constraint-based PC-algorithm is the starting point of many structure learning algorithms, it is focused on finding feature associations by means of conditional independent tests. This is then used to select Bayesian networks, based on a regularized likelihood score. These three topics of data structure analysis were fully tested with synthetic data examples and real cases, which allowed us to unravel and discuss the advantages and limitations of these algorithms. One of the biggest challenges encountered was related to the application of these methods to a Big Data dataset that was analysed within the framework of a collaboration with a large UK retailer, where the interest was in the identification of the data structure underlying customer shopping baskets."],"dc:format":["text"],"dc:identifier.doi":["10.24377/LJMU.t.00010869"],"dc:identifier.uri":["https://researchonline.ljmu.ac.uk/id/eprint/10869/1/2018raulcasanaphd.pdf"],"dc:publisher.department":["Computer Science and Mathematics"],"dc:publisher.institution":["Liverpool John Moores University"],"dc:relation.isreferencedby":["https://researchonline.ljmu.ac.uk/id/eprint/10869/"],"dc:subject":["QA75 Electronic computers. Computer science"],"dc:title":["Identification of Data Structure with Machine Learning: From Fisher to Bayesian networks"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["doctoral"],"dc:type.qualificationname":["phd"]},"updated_at":"2026-07-24T06:30:28Z"}