{"id":{"repo_id":"wku-diss","oai_identifier":"oai:digitalcommons.wku.edu:theses-2531"},"canonical_url":"https://search.dev.ndltd.org/etd/wku-diss/oai:digitalcommons.wku.edu:theses-2531","repository":{"repo_id":"wku-diss","name":"Western Kentucky University","base_url":"https://digitalcommons.wku.edu/do/oai/"},"display":{"title":"An Apache Hadoop Framework for Large-Scale Peptide Identification","abstract":"<p>Peptide identification is an essential step in protein identification, and Peptide Spectrum Match (PSM) data set is huge, which is a time consuming process to work on a single machine. In a typical run of the peptide identification method, PSMs are positioned by a cross correlation, a statistical score, or a likelihood that the match between the trial and hypothetical is correct and unique. This process takes a long time to execute, and there is a demand for an increase in performance to handle large peptide data sets. Development of distributed frameworks are needed to reduce the processing time, but this comes at the price of complexity in developing and executing them. In distributed computing, the program may divide into multiple parts to be executed. The work in this thesis describes the implementation of Apache Hadoop framework for large-scale peptide identification using C-Ranker. The Apache Hadoop data processing software is immersed in a complex environment composed of massive machine clusters, large data sets, and several processing jobs. The framework uses Apache Hadoop Distributed File System (HDFS) and Apache Mapreduce to store and process the peptide data respectively.The proposed framework uses a peptide processing algorithm named CRanker which takes peptide data as an input and identifies the correct PSMs. The framework has two steps: Execute the C-Ranker algorithm on Hadoop cluster and compare the correct PSMs data generated via Hadoop approach with the normal execution approach of C-Ranker. The goal of this framework is to process large peptide datasets using Apache Hadoop distributed approach.</p>","abstract_html":"&lt;p&gt;Peptide identification is an essential step in protein identification, and Peptide Spectrum Match (PSM) data set is huge, which is a time consuming process to work on a single machine. In a typical run of the peptide identification method, PSMs are positioned by a cross correlation, a statistical score, or a likelihood that the match between the trial and hypothetical is correct and unique. This process takes a long time to execute, and there is a demand for an increase in performance to handle large peptide data sets. Development of distributed frameworks are needed to reduce the processing time, but this comes at the price of complexity in developing and executing them. In distributed computing, the program may divide into multiple parts to be executed. The work in this thesis describes the implementation of Apache Hadoop framework for large-scale peptide identification using C-Ranker. The Apache Hadoop data processing software is immersed in a complex environment composed of massive machine clusters, large data sets, and several processing jobs. The framework uses Apache Hadoop Distributed File System (HDFS) and Apache Mapreduce to store and process the peptide data respectively.The proposed framework uses a peptide processing algorithm named CRanker which takes peptide data as an input and identifies the correct PSMs. The framework has two steps: Execute the C-Ranker algorithm on Hadoop cluster and compare the correct PSMs data generated via Hadoop approach with the normal execution approach of C-Ranker. The goal of this framework is to process large peptide datasets using Apache Hadoop distributed approach.&lt;/p&gt;","abstract_has_math":false,"creators":["Donepudi, Harinivesh"],"institution":null,"degree_name":"Master of Science","degree_level":null,"degree_discipline":"Department of Computer Science","degree_department":null,"school":null,"contributors":["Zhonghang Xia (Director), James Gary, Michael Galloway"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2015,"date_issued":"2015-07-01T07:00:00Z","date_published":"2015-07-01T07:00:00Z","updated_at":"2026-07-24T06:08:52Z","subjects":["MapReduce","CRanker","Peptide Spectrum Match","PSM","Biochemistry, Biophysics, and Structural Biology","Computer Sciences","OS and Networks"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.wku.edu/theses/1527","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Zhonghang Xia (Director), James Gary, Michael Galloway"]},{"key":"dc:creator","label":"Author","values":["Donepudi, Harinivesh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Department of Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["MapReduce","CRanker","Peptide Spectrum Match","PSM","Biochemistry, Biophysics, and Structural Biology","Computer Sciences","OS and Networks"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.wku.edu/theses/1527"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Peptide identification is an essential step in protein identification, and Peptide Spectrum Match (PSM) data set is huge, which is a time consuming process to work on a single machine. In a typical run of the peptide identification method, PSMs are positioned by a cross correlation, a statistical score, or a likelihood that the match between the trial and hypothetical is correct and unique. This process takes a long time to execute, and there is a demand for an increase in performance to handle large peptide data sets. Development of distributed frameworks are needed to reduce the processing time, but this comes at the price of complexity in developing and executing them. In distributed computing, the program may divide into multiple parts to be executed. The work in this thesis describes the implementation of Apache Hadoop framework for large-scale peptide identification using C-Ranker. The Apache Hadoop data processing software is immersed in a complex environment composed of massive machine clusters, large data sets, and several processing jobs. The framework uses Apache Hadoop Distributed File System (HDFS) and Apache Mapreduce to store and process the peptide data respectively.The proposed framework uses a peptide processing algorithm named CRanker which takes peptide data as an input and identifies the correct PSMs. The framework has two steps: Execute the C-Ranker algorithm on Hadoop cluster and compare the correct PSMs data generated via Hadoop approach with the normal execution approach of C-Ranker. The goal of this framework is to process large peptide datasets using Apache Hadoop distributed approach.</p>"]},{"key":"dc:title","label":"Title","values":["An Apache Hadoop Framework for Large-Scale Peptide Identification"]}]}],"canonical_facts":{"dc:contributor":["Zhonghang Xia (Director), James Gary, Michael Galloway"],"dc:creator":["Donepudi, Harinivesh"],"dc:description.abstract":["<p>Peptide identification is an essential step in protein identification, and Peptide Spectrum Match (PSM) data set is huge, which is a time consuming process to work on a single machine. In a typical run of the peptide identification method, PSMs are positioned by a cross correlation, a statistical score, or a likelihood that the match between the trial and hypothetical is correct and unique. This process takes a long time to execute, and there is a demand for an increase in performance to handle large peptide data sets. Development of distributed frameworks are needed to reduce the processing time, but this comes at the price of complexity in developing and executing them. In distributed computing, the program may divide into multiple parts to be executed. The work in this thesis describes the implementation of Apache Hadoop framework for large-scale peptide identification using C-Ranker. The Apache Hadoop data processing software is immersed in a complex environment composed of massive machine clusters, large data sets, and several processing jobs. The framework uses Apache Hadoop Distributed File System (HDFS) and Apache Mapreduce to store and process the peptide data respectively.The proposed framework uses a peptide processing algorithm named CRanker which takes peptide data as an input and identifies the correct PSMs. The framework has two steps: Execute the C-Ranker algorithm on Hadoop cluster and compare the correct PSMs data generated via Hadoop approach with the normal execution approach of C-Ranker. The goal of this framework is to process large peptide datasets using Apache Hadoop distributed approach.</p>"],"dc:identifier":["https://digitalcommons.wku.edu/theses/1527"],"dc:subject":["MapReduce","CRanker","Peptide Spectrum Match","PSM","Biochemistry, Biophysics, and Structural Biology","Computer Sciences","OS and Networks"],"dc:title":["An Apache Hadoop Framework for Large-Scale Peptide Identification"],"dc:type":["Thesis"],"thesis:degree_discipline":["Department of Computer Science"],"thesis:degree_name":["Master of Science"]},"updated_at":"2026-07-24T06:08:52Z"}