{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/135554"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/135554","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"A Mutated Peptide Database for the Analysis of Proteomic Mass Spectrometry Data","abstract":"Cancer is often characterized by the accumulation of mutations in a cell population. These mutations often critically disrupt protein function through changes to protein amino acid sequences. Proteomics seeks to characterize and understand proteins through their relative abundance, function, structure, and interactions among others. Proteomic approaches are becoming an increasingly important tool to cancer biologists for profiling cells, tissues, and various healthy or diseased cell states. One of the key technologies that enables protein and peptide amino acid sequences to be determined is mass spectrometry. This is also the premier technology for protein identification. However, it requires the availability of reference protein databases to operate effectively. UniProt's reviewed Swiss-Prot database is one of the best quality protein sequence databases and a valuable tool for protein identifications. This database, however, includes only protein canonical sequences, these being sequences that are widely expressed, functional, and highly confirmed. It contains no reference to mutated sequences. On the other hand, the largest cancer-related mutation repository is the COSMIC database that contains 24,599,940 total variants. The database continues to be updated regularly. In this work, we combined data and information from the well-researched canonical sequence database of Swiss-Prot with the comprehensive list of mutations from the COSMIC database to create a new resource that contains mutated peptides of the human proteome. This process resulted in a new database, termed XMAn, comprising 3,793,617 missense and nonsense mutations. The XMAn v3 database was used to identify mutations in the MDA-MB-231 triple negative breast cancer cell line. A total of 540 unique mutations were identified, of which 60 were also part of the COSMIC's Cancer Gene Census database that incorporates genes implicated in cancer development through their function as oncogenes and tumor suppressor genes, among others. Mutations in the oncogene KRAS, DNA Topoisomerase 1 (TOP1), and TGF-beta receptor type 2 (TGFBR2) represent only a small sample of the mutations that were detected. Overall, the XMAn v3 database proved to be very useful in enabling the identification and characterization of missense and nonsense amino acid level mutations, and providing insights into the possible drivers of aberrant proliferation in the MDA-MB-231 cells. This database will represent a valuable resource to researchers characterizing the proteome of cancer cells, and is aimed to be updated regularly, as COSMIC and UniProt release updates to their respective databases.","abstract_html":"Cancer is often characterized by the accumulation of mutations in a cell population. These mutations often critically disrupt protein function through changes to protein amino acid sequences. Proteomics seeks to characterize and understand proteins through their relative abundance, function, structure, and interactions among others. Proteomic approaches are becoming an increasingly important tool to cancer biologists for profiling cells, tissues, and various healthy or diseased cell states. One of the key technologies that enables protein and peptide amino acid sequences to be determined is mass spectrometry. This is also the premier technology for protein identification. However, it requires the availability of reference protein databases to operate effectively. UniProt&#x27;s reviewed Swiss-Prot database is one of the best quality protein sequence databases and a valuable tool for protein identifications. This database, however, includes only protein canonical sequences, these being sequences that are widely expressed, functional, and highly confirmed. It contains no reference to mutated sequences. On the other hand, the largest cancer-related mutation repository is the COSMIC database that contains 24,599,940 total variants. The database continues to be updated regularly. In this work, we combined data and information from the well-researched canonical sequence database of Swiss-Prot with the comprehensive list of mutations from the COSMIC database to create a new resource that contains mutated peptides of the human proteome. This process resulted in a new database, termed XMAn, comprising 3,793,617 missense and nonsense mutations. The XMAn v3 database was used to identify mutations in the MDA-MB-231 triple negative breast cancer cell line. A total of 540 unique mutations were identified, of which 60 were also part of the COSMIC&#x27;s Cancer Gene Census database that incorporates genes implicated in cancer development through their function as oncogenes and tumor suppressor genes, among others. Mutations in the oncogene KRAS, DNA Topoisomerase 1 (TOP1), and TGF-beta receptor type 2 (TGFBR2) represent only a small sample of the mutations that were detected. Overall, the XMAn v3 database proved to be very useful in enabling the identification and characterization of missense and nonsense amino acid level mutations, and providing insights into the possible drivers of aberrant proliferation in the MDA-MB-231 cells. This database will represent a valuable resource to researchers characterizing the proteome of cancer cells, and is aimed to be updated regularly, as COSMIC and UniProt release updates to their respective databases.","abstract_has_math":false,"creators":["Haueis, Joshua Roman Showalter"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Biological Sciences","degree_department":"Biological Sciences","school":null,"contributors":[],"advisors":[],"committee_chairs":["Lazar, Maria Iuliana"],"committee_members":["Chen, Jing","Suvorov, Anton","Kraikivski, Pavel"],"year":2025,"date_issued":"2025-06-20","date_published":"2025-06-20","updated_at":"2026-07-22T22:19:09Z","subjects":["Mass spectrometry","database","XMAn","mutations","cancer","MDA-231","proteomics","proteins","missense","nonsense"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44144"],"render_values":[{"text":"vt_gsexam:44144","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/135554","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Lazar, Maria Iuliana"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Chen, Jing","Suvorov, Anton","Kraikivski, Pavel"]},{"key":"dc:contributor.department","label":"Department","values":["Biological Sciences"]},{"key":"dc:creator","label":"Author","values":["Haueis, Joshua Roman Showalter"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-06-21T08:01:30Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-06-21T08:01:30Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-06-20"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Biological Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Mass spectrometry","database","XMAn","mutations","cancer","MDA-231","proteomics","proteins","missense","nonsense"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:44144"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/135554"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Cancer is often characterized by the accumulation of mutations in a cell population. These mutations often critically disrupt protein function through changes to protein amino acid sequences. Proteomics seeks to characterize and understand proteins through their relative abundance, function, structure, and interactions among others. Proteomic approaches are becoming an increasingly important tool to cancer biologists for profiling cells, tissues, and various healthy or diseased cell states. One of the key technologies that enables protein and peptide amino acid sequences to be determined is mass spectrometry. This is also the premier technology for protein identification. However, it requires the availability of reference protein databases to operate effectively. UniProt's reviewed Swiss-Prot database is one of the best quality protein sequence databases and a valuable tool for protein identifications. This database, however, includes only protein canonical sequences, these being sequences that are widely expressed, functional, and highly confirmed. It contains no reference to mutated sequences. On the other hand, the largest cancer-related mutation repository is the COSMIC database that contains 24,599,940 total variants. The database continues to be updated regularly. In this work, we combined data and information from the well-researched canonical sequence database of Swiss-Prot with the comprehensive list of mutations from the COSMIC database to create a new resource that contains mutated peptides of the human proteome. This process resulted in a new database, termed XMAn, comprising 3,793,617 missense and nonsense mutations. The XMAn v3 database was used to identify mutations in the MDA-MB-231 triple negative breast cancer cell line. A total of 540 unique mutations were identified, of which 60 were also part of the COSMIC's Cancer Gene Census database that incorporates genes implicated in cancer development through their function as oncogenes and tumor suppressor genes, among others. Mutations in the oncogene KRAS, DNA Topoisomerase 1 (TOP1), and TGF-beta receptor type 2 (TGFBR2) represent only a small sample of the mutations that were detected. Overall, the XMAn v3 database proved to be very useful in enabling the identification and characterization of missense and nonsense amino acid level mutations, and providing insights into the possible drivers of aberrant proliferation in the MDA-MB-231 cells. This database will represent a valuable resource to researchers characterizing the proteome of cancer cells, and is aimed to be updated regularly, as COSMIC and UniProt release updates to their respective databases."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Cancer is a devastating disease that is characterized by the accumulation of mutations. Changes to the DNA of cells leads to changes in the proteins that are expressed in these cells. When these proteins are no longer working as expected, cell populations can become cancerous. Therefore, knowing what mutations have occurred in a population of cells that have lead to cancer development is crucial to understanding the disease. While methods exist to determine the DNA sequences of tumor cells, knowing if their corresponding mutated proteins are being expressed in a cancer cell population is not possible with those methods. The premier technology for gathering protein data from a cell population is mass spectrometry. This technique is able to return the sequences of protein fragments, known as peptides. However, to accurately identify these peptides, reference databases containing protein sequences must be used. One of the best databases to do this is known as UniProt, more specifically the highly reviewed Swiss-Prot database. This database, unfortunately, does not contain a robust concentration of mutated peptides. The COSMIC database is the most robust dataset containing mutations identified in cancer cells. However, it does not contain the protein sequence data that Swiss-Prot provides. Therefore, the aim of this work was to generate a mutated database, XMAn v3, which combines the robust sequence data of Swiss-Prot with the large quantity of mutation data in COSMIC, to give us a 3,793,617 entry reference database for use in mass spectrometry. To demonstrate the utility of the newly created database, an analysis of the MDA-MB-231 triple negative breast cancer cell line was conducted. A total of 540 unique mutations were identified, with interesting mutations in genes such as KRAS, DNA Topoisomerase 1 (TOP1), and TGF-beta receptor type 2 (TGFBR2). Overall, the XMAn v3 database proved to be very useful in identifying mutated peptides. This database will represent a valuable resource to researchers searching for mutation spots in cancer cells, and is aimed to be updated regularly, as COSMIC and UniProt will continue to release updates to their respective databases."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["A Mutated Peptide Database for the Analysis of Proteomic Mass Spectrometry Data"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Lazar, Maria Iuliana"],"dc:contributor.committeemember":["Chen, Jing","Suvorov, Anton","Kraikivski, Pavel"],"dc:contributor.department":["Biological Sciences"],"dc:creator":["Haueis, Joshua Roman Showalter"],"dc:date.accessioned":["2025-06-21T08:01:30Z"],"dc:date.available":["2025-06-21T08:01:30Z"],"dc:date.issued":["2025-06-20"],"dc:description.abstract":["Cancer is often characterized by the accumulation of mutations in a cell population. These mutations often critically disrupt protein function through changes to protein amino acid sequences. Proteomics seeks to characterize and understand proteins through their relative abundance, function, structure, and interactions among others. Proteomic approaches are becoming an increasingly important tool to cancer biologists for profiling cells, tissues, and various healthy or diseased cell states. One of the key technologies that enables protein and peptide amino acid sequences to be determined is mass spectrometry. This is also the premier technology for protein identification. However, it requires the availability of reference protein databases to operate effectively. UniProt's reviewed Swiss-Prot database is one of the best quality protein sequence databases and a valuable tool for protein identifications. This database, however, includes only protein canonical sequences, these being sequences that are widely expressed, functional, and highly confirmed. It contains no reference to mutated sequences. On the other hand, the largest cancer-related mutation repository is the COSMIC database that contains 24,599,940 total variants. The database continues to be updated regularly. In this work, we combined data and information from the well-researched canonical sequence database of Swiss-Prot with the comprehensive list of mutations from the COSMIC database to create a new resource that contains mutated peptides of the human proteome. This process resulted in a new database, termed XMAn, comprising 3,793,617 missense and nonsense mutations. The XMAn v3 database was used to identify mutations in the MDA-MB-231 triple negative breast cancer cell line. A total of 540 unique mutations were identified, of which 60 were also part of the COSMIC's Cancer Gene Census database that incorporates genes implicated in cancer development through their function as oncogenes and tumor suppressor genes, among others. Mutations in the oncogene KRAS, DNA Topoisomerase 1 (TOP1), and TGF-beta receptor type 2 (TGFBR2) represent only a small sample of the mutations that were detected. Overall, the XMAn v3 database proved to be very useful in enabling the identification and characterization of missense and nonsense amino acid level mutations, and providing insights into the possible drivers of aberrant proliferation in the MDA-MB-231 cells. This database will represent a valuable resource to researchers characterizing the proteome of cancer cells, and is aimed to be updated regularly, as COSMIC and UniProt release updates to their respective databases."],"dc:description.abstractgeneral":["Cancer is a devastating disease that is characterized by the accumulation of mutations. Changes to the DNA of cells leads to changes in the proteins that are expressed in these cells. When these proteins are no longer working as expected, cell populations can become cancerous. Therefore, knowing what mutations have occurred in a population of cells that have lead to cancer development is crucial to understanding the disease. While methods exist to determine the DNA sequences of tumor cells, knowing if their corresponding mutated proteins are being expressed in a cancer cell population is not possible with those methods. The premier technology for gathering protein data from a cell population is mass spectrometry. This technique is able to return the sequences of protein fragments, known as peptides. However, to accurately identify these peptides, reference databases containing protein sequences must be used. One of the best databases to do this is known as UniProt, more specifically the highly reviewed Swiss-Prot database. This database, unfortunately, does not contain a robust concentration of mutated peptides. The COSMIC database is the most robust dataset containing mutations identified in cancer cells. However, it does not contain the protein sequence data that Swiss-Prot provides. Therefore, the aim of this work was to generate a mutated database, XMAn v3, which combines the robust sequence data of Swiss-Prot with the large quantity of mutation data in COSMIC, to give us a 3,793,617 entry reference database for use in mass spectrometry. To demonstrate the utility of the newly created database, an analysis of the MDA-MB-231 triple negative breast cancer cell line was conducted. A total of 540 unique mutations were identified, with interesting mutations in genes such as KRAS, DNA Topoisomerase 1 (TOP1), and TGF-beta receptor type 2 (TGFBR2). Overall, the XMAn v3 database proved to be very useful in identifying mutated peptides. This database will represent a valuable resource to researchers searching for mutation spots in cancer cells, and is aimed to be updated regularly, as COSMIC and UniProt will continue to release updates to their respective databases."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:44144"],"dc:identifier.uri":["https://hdl.handle.net/10919/135554"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Mass spectrometry","database","XMAn","mutations","cancer","MDA-231","proteomics","proteins","missense","nonsense"],"dc:title":["A Mutated Peptide Database for the Analysis of Proteomic Mass Spectrometry Data"],"dc:type":["Thesis"],"thesis:degree_discipline":["Biological Sciences"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:19:09Z"}