{"id":{"repo_id":"calpoly","oai_identifier":"oai:digitalcommons.calpoly.edu:theses-1583"},"canonical_url":"https://search.dev.ndltd.org/etd/calpoly/oai:digitalcommons.calpoly.edu:theses-1583","repository":{"repo_id":"calpoly","name":"Cal Poly","base_url":"https://digitalcommons.calpoly.edu/do/oai/"},"display":{"title":"An Efficient Primer Selection Process Combining Progressive and Iterative Multiple Sequence Alignment Strategies: ClustalW and HMMER","abstract":"<p> <p>This thesis describes a method for using a computationally efficient algorithm to identify candidate DNA primer sequences. DNA sequencing primers are a critical element of polymerase chain reaction (PCR) and DNA sequence analysis. A variety of methods for deriving DNA primers exist, but such methods are often computationally intensive, or do not use available sequence data that could potentially serve as a possible resource for primer identification. Though no current algorithm exists which will always yield a correct primer for every need, evaluation of multi-sequence alignments may provide a reliable source for primer candidates. However, an exact mathematical solution for multi-sequence alignments, using currently available computational resources, is only viable for a very small number of sequences. Any solution for a larger number of sequences will therefore use other computational methods and heuristics to estimate an alignment.</p> <p>The solution presented here, featuring a combination of ClustalW and HMMER alignment tools, is able to identify conserved regions in sequence data in a computationally efficient manner, and from these regions, suggest viable primer candidates. Computational complexity for the HMMER alignment effort has been maintained at O(MN); the suggested process for creating sequence alignments lead to a 15-fold improvement in performance over conventional methods, while also successfully identifying fungal specific primers, with individual examples showing 90% or greater match for the given fungal phylum.</p> <p>It was found that alignment quality could be further improved by using simple sorting methods against input sequence data.</p> </p>","abstract_html":"&lt;p&gt; &lt;p&gt;This thesis describes a method for using a computationally efficient algorithm to identify candidate DNA primer sequences. DNA sequencing primers are a critical element of polymerase chain reaction (PCR) and DNA sequence analysis. A variety of methods for deriving DNA primers exist, but such methods are often computationally intensive, or do not use available sequence data that could potentially serve as a possible resource for primer identification. Though no current algorithm exists which will always yield a correct primer for every need, evaluation of multi-sequence alignments may provide a reliable source for primer candidates. However, an exact mathematical solution for multi-sequence alignments, using currently available computational resources, is only viable for a very small number of sequences. Any solution for a larger number of sequences will therefore use other computational methods and heuristics to estimate an alignment.&lt;/p&gt; &lt;p&gt;The solution presented here, featuring a combination of ClustalW and HMMER alignment tools, is able to identify conserved regions in sequence data in a computationally efficient manner, and from these regions, suggest viable primer candidates. Computational complexity for the HMMER alignment effort has been maintained at O(MN); the suggested process for creating sequence alignments lead to a 15-fold improvement in performance over conventional methods, while also successfully identifying fungal specific primers, with individual examples showing 90% or greater match for the given fungal phylum.&lt;/p&gt; &lt;p&gt;It was found that alignment quality could be further improved by using simple sorting methods against input sequence data.&lt;/p&gt; &lt;/p&gt;","abstract_has_math":false,"creators":["Green, Michael C."],"institution":null,"degree_name":"MS in Computer Science","degree_level":null,"degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Timothy J. Kearns"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-06-01T07:00:00Z","date_published":"2011-06-01T07:00:00Z","updated_at":"2026-07-24T01:32:55Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["10.15368/theses.2011.104"],"render_values":[{"text":"10.15368/theses.2011.104","href":"https://doi.org/10.15368/theses.2011.104","code":true}]}]},"links":{"outbound_url":"https://digitalcommons.calpoly.edu/theses/547","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Timothy J. Kearns"]},{"key":"dc:creator","label":"Author","values":["Green, Michael C."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2011-06-14T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_name","label":"Degree Name","values":["MS in Computer Science"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.calpoly.edu/theses/547","10.15368/theses.2011.104"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p> <p>This thesis describes a method for using a computationally efficient algorithm to identify candidate DNA primer sequences. DNA sequencing primers are a critical element of polymerase chain reaction (PCR) and DNA sequence analysis. A variety of methods for deriving DNA primers exist, but such methods are often computationally intensive, or do not use available sequence data that could potentially serve as a possible resource for primer identification. Though no current algorithm exists which will always yield a correct primer for every need, evaluation of multi-sequence alignments may provide a reliable source for primer candidates. However, an exact mathematical solution for multi-sequence alignments, using currently available computational resources, is only viable for a very small number of sequences. Any solution for a larger number of sequences will therefore use other computational methods and heuristics to estimate an alignment.</p> <p>The solution presented here, featuring a combination of ClustalW and HMMER alignment tools, is able to identify conserved regions in sequence data in a computationally efficient manner, and from these regions, suggest viable primer candidates. Computational complexity for the HMMER alignment effort has been maintained at O(MN); the suggested process for creating sequence alignments lead to a 15-fold improvement in performance over conventional methods, while also successfully identifying fungal specific primers, with individual examples showing 90% or greater match for the given fungal phylum.</p> <p>It was found that alignment quality could be further improved by using simple sorting methods against input sequence data.</p> </p>"]},{"key":"dc:title","label":"Title","values":["An Efficient Primer Selection Process Combining Progressive and Iterative Multiple Sequence Alignment Strategies: ClustalW and HMMER"]}]}],"canonical_facts":{"dc:contributor":["Timothy J. Kearns"],"dc:creator":["Green, Michael C."],"dc:date.available":["2011-06-14T07:00:00Z"],"dc:description.abstract":["<p> <p>This thesis describes a method for using a computationally efficient algorithm to identify candidate DNA primer sequences. DNA sequencing primers are a critical element of polymerase chain reaction (PCR) and DNA sequence analysis. A variety of methods for deriving DNA primers exist, but such methods are often computationally intensive, or do not use available sequence data that could potentially serve as a possible resource for primer identification. Though no current algorithm exists which will always yield a correct primer for every need, evaluation of multi-sequence alignments may provide a reliable source for primer candidates. However, an exact mathematical solution for multi-sequence alignments, using currently available computational resources, is only viable for a very small number of sequences. Any solution for a larger number of sequences will therefore use other computational methods and heuristics to estimate an alignment.</p> <p>The solution presented here, featuring a combination of ClustalW and HMMER alignment tools, is able to identify conserved regions in sequence data in a computationally efficient manner, and from these regions, suggest viable primer candidates. Computational complexity for the HMMER alignment effort has been maintained at O(MN); the suggested process for creating sequence alignments lead to a 15-fold improvement in performance over conventional methods, while also successfully identifying fungal specific primers, with individual examples showing 90% or greater match for the given fungal phylum.</p> <p>It was found that alignment quality could be further improved by using simple sorting methods against input sequence data.</p> </p>"],"dc:identifier":["https://digitalcommons.calpoly.edu/theses/547","10.15368/theses.2011.104"],"dc:title":["An Efficient Primer Selection Process Combining Progressive and Iterative Multiple Sequence Alignment Strategies: ClustalW and HMMER"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_name":["MS in Computer Science"]},"updated_at":"2026-07-24T01:32:55Z"}