Back to results

Technische Universität Berlin

Computational methods and machine learning for crosslinking mass spectrometry data analysis

Abstract

dc:description.abstract

A central part in understanding complex biological systems is to uncover the function and structure of proteins. The elucidation of a protein’s structure and understanding its function are tightly connected. The underlying paradigm that structure defines function, has led to the development of many methods to derive the three-dimensional structure of proteins and protein complexes. Crosslinking mass spectrometry (CLMS) is a comparatively new tool for the analysis of single proteins, multi-protein complexes, and protein-protein interactions. CLMS poses several challenges for mass spectrometry-based proteomics, which include understanding the fragmentation behavior of crosslinked peptides to design efficient database search strategies and improved acquisition settings. CLMS builds upon the preservation of distance information by crosslinking reagents, which is relayed by mass spectrometric analysis. To identify a crosslink in a standard database search, theoretically all pairwise peptide combinations need to be considered. Without the use of isotope-labeled or cleavable crosslinkers, applying a standard crosslinking approach using homobifunctional NHS-ester crosslinker reagents, an exhaustive peptide identification strategy becomes quickly unfeasible because of the dynamic explosion of the search space. Therefore, robust heuristics are needed to make the identification of crosslinks in complex samples feasible. This endeavor is even further hindered by the unequal fragmentation of the two peptides in a crosslink under collision-induced dissociation conditions. The subsequent coverage gap between the two peptides in a crosslink may lead to misidentifications. This thesis presents computational approaches and machine learning methods to improve the identification of crosslinked peptides. First, an efficient strategy is outlined, based on an explorative study about the fragmentation behavior of crosslinked peptides. Most importantly, the presented search strategy shows that the information from isotope-labeled and cleavable crosslinkers can be partially retrieved by computational processing of the spectra and adequate mass spectrometric acquisition settings. A key concept builds upon the ability to recognize crosslinked fragments from their mass and charge. This allows to identify the two linked peptides in a sequential manner without searching all peptide combinations exhaustively. Second, to reduce the coverage gap, modern mass spectrometers offer versatile fragmentation methods. For most crosslinks, electron-transfer dissociation combined with higher-energy collision dissociation (HCD) yields the highest coverage.HCD remains an important choice because of its fast acquisition speed and competitive sequence coverage. Third, to avoid severe bias through the identification of noncovalently associated peptides as crosslinks, multiple solutions are feasible. For example, disruptive ionization settings can be used to avoid noncovalently associated peptides entering the mass spectrometer. Alternatively, post-acquisition heuristics using the retention time difference between linear and crosslinked peptides add valuable information to recognize noncovalent peptide associations. Fourth, since complex crosslinking experiments with deep-proteome coverage require extensive fractionation, being able to predict the retention behavior may prove beneficial for peptide identification. In addition, mechanistic understanding of the separation process helps to further improve the chromatographic separation. For hydrophilic anion exchange chromatography (hSAX), the separation is heavily influenced by charged amino acids and aromatics. Most importantly, the retention behavior of linear peptides can be accurately predicted through deep neural networks. Fifth, the ability to predict not only hSAX, but also strong cation exchange (SCX) and reversed-phase retention times indeed proves to be a valuable addition for the identification of crosslinked peptides. Siamese neural network architectures offer elegant solutions to encode crosslinked peptides. Multi-task learning of several chromatography domains at the same time allows robust and fast prediction of all chromatography domains. Accurate reversed-phase predictions together with hSAX and SCX fraction prediction allows rescoring already identified peptide spectrum matches with a support vector machine. This workflow leads to more identified protein-protein interactions at constant false discovery rate from a deep-fractionated Escherichia coli sample. The integration of advancements in crosslinking chemistry, sample acquisition, database search, and machine learning together are essential stepping-stones for the identification of crosslinked peptides in complex samples.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Giese, Sven Hans-Joachim
Advisor dc:contributor.advisor
  • Rappsilber, Juri

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
OAI identifier oai:identifier
oai:depositonce.tu-berlin.de:11303/13236

Chain of custody

source
Harvested from
Technische Universität Berlin
Base URL
api-depositonce.tu-berlin.de/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
related terms
citation

Giese, Sven Hans-Joachim. Computational methods and machine learning for crosslinking mass spectrometry data analysis. 2021. https://depositonce.tu-berlin.de/handle/11303/13236