Back to search

University of British Columbia

Data analysis in proteomics novel computational strategies for modeling and interpreting complex mass spectrometry data

Abstract

dc:description

Contemporary proteomics studies require computational approaches to deal with both the complexity of the data generated, and with the volume of data produced. The amalgamation of mass spectrometry -- the analytical tool of choice in proteomics -- with the computational and statistical sciences is still recent, and several avenues of exploratory data analysis and statistical methodology remain relatively unexplored. The current study focuses on three broad analytical domains, and develops novel exploratory approaches and practical tools in each. Data transform approaches are the first explored. These methods re-frame data, allowing for the visualization and exploitation of features and trends that are not immediately evident. An exploratory approach making use of the correlation transform is developed, and is used to identify mass-shift signals in mass spectra. This approach is used to identify and map post-translational modifications on individual peptides, and to identify SILAC modification-containing spectra in a full-scale proteomic analysis. Secondly, matrix decomposition and projection approaches are explored; these use an eigen-decomposition to extract general trends from groups of related spectra. A data visualization approach is demonstrated using these techniques, capable of visualizing trends in large numbers of complex spectra, and a data compression and feature extraction technique is developed suitable for use in spectral modeling. Finally, a general machine learning approach is developed based on conditional random fields (CRFs). These models are capable of dealing with arbitrary sequence modeling tasks, similar to hidden Markov models (HMMs), but are far more robust to interdependent observational features, and do not require limiting independence assumptions to remain tractable. The theory behind this approach is developed, and a simple machine learning fragmentation model is developed to test the hypothesis that reproducible sequence-specific intensity ratios are present within the distribution of fragment ions originating from a common peptide bond breakage. After training, the model shows very good performance associating peptide sequences and fragment ion intensity information, lending strong support to the hypothesis.

Degree

thesis:*
Name thesis:degree_name
Master of Science - MSc
Level thesis:degree_level
master's
Discipline thesis:degree_discipline
Experimental Medicine
Grantor dc:publisher
University of British Columbia
Year dc:date
2008

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sniatynski, Matthew John

Rights

dc:rights
Statement dc:rights
  • Attribution-NonCommercial-NoDerivatives 4.0 International
Language dc:language
eng

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2429/1497
OAI identifier oai:identifier
oai:circle.library.ubc.ca:2429/1497

Chain of custody

source
Harvested from
University of British Columbia
Base URL
circle.library.ubc.ca/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Sniatynski, Matthew John. Data analysis in proteomics novel computational strategies for modeling and interpreting complex mass spectrometry data. master's thesis, University of British Columbia, 2008. http://hdl.handle.net/2429/1497