Back to results

Massachusetts Institute of Technology

Machine Learning Methods for High Throughput Biological Data

Abstract

dc:description.abstract

Machine learning is becoming a pivotal tool in the analysis of datasets generated from high-throughput biological omics experiments. However, omics data introduces distinctive algorithmic challenges that set it apart from other domains where machine learning is applied. These challenges encompass issues such as limited data availability, complex noise, ambiguities in representation, and the absence of definitive ground truth for validation. In this thesis, I present three examples of machine learning applications to different omics modalities in which I address these challenges. In my first project, I develop an approach for contrastive representation learning with immunohistochemistry images, which suffer complex technical and biological noise that render generic approaches ineffective; and I demonstrate how this approach can be combined with noisy labels derived from transcriptomics to derive an effective classifier of cell-type specificity. In my second project, I consider the problem of predicting mass spectra of small molecules: previous methods suffer from a tradeoff between capturing high-resolution mass information and a tractable learning problem, which I resolve by introducing a novel representation of the output space. In my third project, I perform gene regulatory network inference using a number of different single-cell sequencing platforms, and carry out a quantitative comparison of these technologies. In summary, this thesis showcases the difficulties that arise in applying modern machine learning approaches to high-throughput biological measurements, and empirical case studies of how these difficulties may be overcome.

Degree

thesis:*
Name thesis:degree_name
Doctoral
Department dc:contributor.department
Massachusetts Institute of Technology. Computational and Systems Biology Program
Grantor dc:publisher
Massachusetts Institute of Technology
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Murphy, Michael A.
Advisors dc:contributor.advisor
  • Fraenkel, Ernest
  • Jegelka, Stefanie

Rights

dc:rights
Statement dc:rights
  • Attribution 4.0 International (CC BY 4.0)
  • Copyright retained by author(s)

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/1721.1/154024
OAI identifier oai:identifier
oai:dspace.mit.edu:1721.1/154024

Chain of custody

source
Harvested from
MIT
Base URL
dspace.mit.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
related terms
citation

Murphy, Michael A.. Machine Learning Methods for High Throughput Biological Data. Massachusetts Institute of Technology, 2024. https://hdl.handle.net/1721.1/154024