{"id":{"repo_id":"wfu","oai_identifier":"oai:wakespace.lib.wfu.edu:10339/90706"},"canonical_url":"https://search.dev.ndltd.org/etd/wfu/oai:wakespace.lib.wfu.edu:10339/90706","repository":{"repo_id":"wfu","name":"Wake Forest University","base_url":"https://wakespace.lib.wfu.edu/oai/request"},"display":{"title":"CLASSIFYING PEROXIREDOXIN SUBGROUPS AND IDENTIFYING DISCRIMINATING MOTIFS VIA MACHINE LEARNING","abstract":"Accurate and automated functional annotation is a pressing open problem, with functional characterizations lagging far behind the exponential growth in biological sequence databases. In this thesis, I present our recent development of machine learning methods for high-throughput, accurate, sequence-based functional annotation. Chapter 1 describes the biological and computational background of this study. Chapter 2defines the specific problem we try to solve. Chapter 3 demonstrates that our 3mer-SVM, that accurately classifies Peroxiredoxin subgroups, can provide meaningful additional insight into the functional conserved sites in Peroxiredoxin protein. Moreover, in Chapter 4, we propose a two-round learning algorithm that can capture gapped-kmer features in sequences and lead to more accurate classifications than the kmer-SVM approach. We illustrate this learning algorithm can be useful as a de novo motif finder for uncovering discriminating motifs among sequences associated with particular activities and functions. With a brief discussion on the advantage and limitations on our kmer-based sequence classification and \\textit{de novo} motif identification, in Chapter 5, we propose several potential applications for future directions.","abstract_html":"Accurate and automated functional annotation is a pressing open problem, with functional characterizations lagging far behind the exponential growth in biological sequence databases. In this thesis, I present our recent development of machine learning methods for high-throughput, accurate, sequence-based functional annotation. Chapter 1 describes the biological and computational background of this study. Chapter 2defines the specific problem we try to solve. Chapter 3 demonstrates that our 3mer-SVM, that accurately classifies Peroxiredoxin subgroups, can provide meaningful additional insight into the functional conserved sites in Peroxiredoxin protein. Moreover, in Chapter 4, we propose a two-round learning algorithm that can capture gapped-kmer features in sequences and lead to more accurate classifications than the kmer-SVM approach. We illustrate this learning algorithm can be useful as a de novo motif finder for uncovering discriminating motifs among sequences associated with particular activities and functions. With a brief discussion on the advantage and limitations on our kmer-based sequence classification and \\textit{de novo} motif identification, in Chapter 5, we propose several potential applications for future directions.","abstract_has_math":false,"creators":["Xiao, Jiajie"],"institution":"Wake Forest University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018","date_published":"2018","updated_at":"2026-07-27T22:02:17Z","subjects":[],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/10339/90706","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Xiao, Jiajie"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2018-05-24T08:36:02Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2019-05-23T08:30:11Z"]},{"key":"dc:date.issued","label":"Date","values":["2018"]},{"key":"dc:publisher","label":"Institution","values":["Wake Forest University"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/10339/90706"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Accurate and automated functional annotation is a pressing open problem, with functional characterizations lagging far behind the exponential growth in biological sequence databases. In this thesis, I present our recent development of machine learning methods for high-throughput, accurate, sequence-based functional annotation. Chapter 1 describes the biological and computational background of this study. Chapter 2defines the specific problem we try to solve. Chapter 3 demonstrates that our 3mer-SVM, that accurately classifies Peroxiredoxin subgroups, can provide meaningful additional insight into the functional conserved sites in Peroxiredoxin protein. Moreover, in Chapter 4, we propose a two-round learning algorithm that can capture gapped-kmer features in sequences and lead to more accurate classifications than the kmer-SVM approach. We illustrate this learning algorithm can be useful as a de novo motif finder for uncovering discriminating motifs among sequences associated with particular activities and functions. With a brief discussion on the advantage and limitations on our kmer-based sequence classification and \\textit{de novo} motif identification, in Chapter 5, we propose several potential applications for future directions."]},{"key":"dc:title","label":"Title","values":["CLASSIFYING PEROXIREDOXIN SUBGROUPS AND IDENTIFYING DISCRIMINATING MOTIFS VIA MACHINE LEARNING"]}]}],"canonical_facts":{"dc:creator":["Xiao, Jiajie"],"dc:date.accessioned":["2018-05-24T08:36:02Z"],"dc:date.available":["2019-05-23T08:30:11Z"],"dc:date.issued":["2018"],"dc:description.abstract":["Accurate and automated functional annotation is a pressing open problem, with functional characterizations lagging far behind the exponential growth in biological sequence databases. In this thesis, I present our recent development of machine learning methods for high-throughput, accurate, sequence-based functional annotation. Chapter 1 describes the biological and computational background of this study. Chapter 2defines the specific problem we try to solve. Chapter 3 demonstrates that our 3mer-SVM, that accurately classifies Peroxiredoxin subgroups, can provide meaningful additional insight into the functional conserved sites in Peroxiredoxin protein. Moreover, in Chapter 4, we propose a two-round learning algorithm that can capture gapped-kmer features in sequences and lead to more accurate classifications than the kmer-SVM approach. We illustrate this learning algorithm can be useful as a de novo motif finder for uncovering discriminating motifs among sequences associated with particular activities and functions. With a brief discussion on the advantage and limitations on our kmer-based sequence classification and \\textit{de novo} motif identification, in Chapter 5, we propose several potential applications for future directions."],"dc:identifier.uri":["http://hdl.handle.net/10339/90706"],"dc:language.iso":["en"],"dc:publisher":["Wake Forest University"],"dc:title":["CLASSIFYING PEROXIREDOXIN SUBGROUPS AND IDENTIFYING DISCRIMINATING MOTIFS VIA MACHINE LEARNING"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T22:02:17Z"}