Back to results

Liverpool John Moores University

Fisher networks: A principled approach to retrieval-based classification

Abstract

dc:description.abstract

Due to the technological advances in the acquisition and processing of information, current data mining applications involve databases of sizes that would be unthinkable just two decades ago. However, real-word datasets are often riddled with irrelevant variables that not only do not generate any meaningful information about the process of interest, but may also obstruct the contribution of the truly informative data features. Taking into consideration the relevance of the different measures available can make the difference between reaching an accurate reflection of the underlying truth and obtaining misleading results that cause the drawing of erroneousconclusions. Another important consideration in data analysis is the interpretability of the models used to fit the data. It is clear that performance must be a key aspect in deciding which methodology to use, but it should not be the only one. Models with an obscure internal operation see their practical usefulness effectively diminished by the difficulty to understand the reasoning behind their inferences, which makes them less appealing to users that are not familiar with their theoretical basis. This thesis proposes a novel framework for the visualisation and categorisation of data in classification contexts that tackles the two issues discussed above and provides an informative output of intuitive interpretation. The system is based on a Fisher information metric that automatically filters the contribution of variables depending on their relevance with respect to the classification problem at hand, measured by their influence on the posterior class probabilities. Fisher distances can then be used to calculate rigorous problem-specific similarity measures, which can be grouped into a pairwise adjacency matrix, thus defining a network. Following this novel construction process results in a principled visualisation of the data organised in communities that highlights the structure of the underlying class membership probabilities. Furthermore, the relational nature of the network can be used to reproduce the probabilistic predictions of the original estimates in a case-based approach, making them explainable by means of known cases in the dataset. The potential applications and usefulness of the framework are illustrated using several real-world datasets, giving examples of the typical output that the end user receives and how they can use it to learn more about the cases of interest as well as about the dataset as a whole.

Degree

thesis:*
Name dc:type.qualificationname
phd
Level dc:type.qualificationlevel
doctoral
Grantor dc:publisher.institution
Liverpool John Moores University
Year dc:date.issued
2013

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ruiz, H
Contributors dc:contributor
  • Lisboa, P
  • Jarman, I
  • Martin-Guerrero, JD

Subjects

dc:subject × 2

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:researchonline.ljmu.ac.uk:4371

Chain of custody

source
Harvested from
Liverpool Jon Moores University
Base URL
researchonline.ljmu.ac.uk/cgi/oai2
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Ruiz, H. Fisher networks: A principled approach to retrieval-based classification. doctoral thesis, Liverpool John Moores University, 2013. https://doi.org/10.24377/LJMU.t.00004371