Massachusetts Institute of Technology
Using Co-evolutionary Information to Improve Protein Language Modelling
Abstract
dc:description.abstractProtein engineering has the potential to solve complex global problems in medicine, clean energy, and manufacturing. However, current protein engineering efforts are hampered by a lack of supervised data. We help recitify this issue by developing supervised models that perform well in data-constrained settings by generalizing across protein engineering tasks and better incorporating coevolutionary and structural information. We also develop an unsupervised language model that conditions the target sequence on its multiple sequence alignment, allowing us to better model protein families.
Degree
thesis:*- Name thesis:degree_name
- Master
- Department dc:contributor.department
- Massachusetts Institute of Technology. Department of Electrical Engineering and Computer Science
- Grantor dc:publisher
- Massachusetts Institute of Technology
- Year dc:date.issued
- 2021
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ram, Soumya
- Advisor dc:contributor.advisor
-
- Bepler, Tristan
Rights
dc:rights- Statement dc:rights
-
- In Copyright - Educational Use Permitted
- Copyright MIT
- Licence dc:rights.uri
Identifiers
dc:identifier.*- Handle dc:identifier.uri
- https://hdl.handle.net/1721.1/139337
- OAI identifier oai:identifier
- oai:dspace.mit.edu:1721.1/139337