Back to results

Massachusetts Institute of Technology

Machine learning for understanding protein sequence and structure

Abstract

dc:description.abstract

Proteins are the fundamental building blocks of life, carrying out a vast array of functions at the molecular level. Understanding these molecular machines has been a core problem in biology for decades. Recent advances in cryo-electron microscopy (cryoEM) has enabled high resolution experimental measurement of proteins in their native states. However, this technology remains expensive and low throughput. At the same time, ever growing protein databases offer new opportunities for understanding the diversity of natural proteins and for linking sequence to structure and function. This thesis introduces a variety of machine learning methods for accelerating protein structure determination by cryoEM and for learning from large protein databases. We first consider the problem of protein identification in the large images collected in cryoEM. We propose a positive-unlabeled learning framework that enables high accuracy particle detection with few labeled data points, both improving data quality and analysis speed. Next, we develop a deep denoising model for cryo-electron micrographs. By learning the denoising model from large amounts of real cryoEM data, we are able to capture the noise generation process and accurately denoise micrographs, improving the ability of experamentalists to examine and interpret their data. We then introduce a neural network model for understanding continuous variability in proteins in cryoEM data by explicitly disentangling variation of interest (structure) for nuisance variation due to rotation and translation. Finally, we move beyond cryoEM and propose a method for learning vector embeddings of proteins using information from structure and sequence. Many of the machine learning methods developed here are general purpose and can be applied to other data domains.

Degree

thesis:*
Name thesis:degree_name
Doctoral
Department dc:contributor.department
Massachusetts Institute of Technology. Computational and Systems Biology Program
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bepler, Tristan(Tristan Wendland)
Advisor dc:contributor.advisor
  • Bonnie Berger.

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • MIT theses may be protected by copyright. Please reuse MIT thesis content according to the MIT Libraries Permissions Policy, which is available through the URL provided.
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1721.1/129888
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/129888

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Bepler, Tristan(Tristan Wendland). Machine learning for understanding protein sequence and structure. Massachusetts Institute of Technology, 2020. https://hdl.handle.net/1721.1/129888