{"id":{"repo_id":"unr","oai_identifier":"oai:scholarwolf.unr.edu:11714/11430"},"canonical_url":"https://search.dev.ndltd.org/etd/unr/oai:scholarwolf.unr.edu:11714/11430","repository":{"repo_id":"unr","name":"University of Nevada - Reno","base_url":"https://scholarwolf.unr.edu/server/oai/request"},"display":{"title":"Decoding DNA: scalable genome annotation software, generative AI, and application to the cactus pear genome","abstract":"The annotation of protein-coding genes from a raw genome assembly involves identifying the sequence regions that are transcribed into mRNA transcripts, spliced, and ultimately translated into protein. Thus, genome annotation is a fundamental task that critically supports systems level biological inquiry by providing a reference for 'omics' analyses including, RNA sequencing, proteomics, and comparative genomics. Despite advances in genome sequencing, which enable routine chromosome-level genome assembly, complete identification of gene structures in eukaryotes remains a significant challenge. For example, existing genome annotation pipelines suffer from a lack of automation, an inability to control precision, and from poor performance in predicting alternative splicing.The goal of this work is to develop improved computational tools for genome annotation and to apply the improved methods to annotate the genome of cactus pear (Opuntia cochenillifera). To address challenges in computational efficiency and precision, an automated bioinformatics pipeline called Sylvan was developed that computes a comprehensive genome annotation from disparate evidence sources and filters spurious gene models using a semi-supervised random forest classifier. In benchmarking trials involving Arabidopsis thaliana and Oryza sativa the pipeline outperformed current standards, such as MAKER and BRAKER, in both F1 similarity and BUSCO completeness. Sylvan was used to annotate the genome of Opuntia cochenillifera, representing the first genome sequence and assembly in the genus and a foundational tool for research into crassulacean acid metabolism (CAM) and drought tolerance in plants. To improve the capacity of genome annotation tools to predict full-length, alternatively spliced transcripts ab initio, a deep learning transformer model was developed to 'translate' a DNA sequence into its text-based annotation. This generative strategy provides increased flexibility to predict hierarchical and overlapping gene structures that are not possible with one dimensional segmentation models.","abstract_html":"The annotation of protein-coding genes from a raw genome assembly involves identifying the sequence regions that are transcribed into mRNA transcripts, spliced, and ultimately translated into protein. Thus, genome annotation is a fundamental task that critically supports systems level biological inquiry by providing a reference for &#x27;omics&#x27; analyses including, RNA sequencing, proteomics, and comparative genomics. Despite advances in genome sequencing, which enable routine chromosome-level genome assembly, complete identification of gene structures in eukaryotes remains a significant challenge. For example, existing genome annotation pipelines suffer from a lack of automation, an inability to control precision, and from poor performance in predicting alternative splicing.The goal of this work is to develop improved computational tools for genome annotation and to apply the improved methods to annotate the genome of cactus pear (Opuntia cochenillifera). To address challenges in computational efficiency and precision, an automated bioinformatics pipeline called Sylvan was developed that computes a comprehensive genome annotation from disparate evidence sources and filters spurious gene models using a semi-supervised random forest classifier. In benchmarking trials involving Arabidopsis thaliana and Oryza sativa the pipeline outperformed current standards, such as MAKER and BRAKER, in both F1 similarity and BUSCO completeness. Sylvan was used to annotate the genome of Opuntia cochenillifera, representing the first genome sequence and assembly in the genus and a foundational tool for research into crassulacean acid metabolism (CAM) and drought tolerance in plants. To improve the capacity of genome annotation tools to predict full-length, alternatively spliced transcripts ab initio, a deep learning transformer model was developed to &#x27;translate&#x27; a DNA sequence into its text-based annotation. This generative strategy provides increased flexibility to predict hierarchical and overlapping gene structures that are not possible with one dimensional segmentation models.","abstract_has_math":false,"creators":["Lomas, Johnathan"],"institution":null,"degree_name":null,"degree_level":"Doctorate Degree","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Cushman, John C.","Yim, Won C."],"committee_chairs":[],"committee_members":["Harper, Jeffrey F.","Choi, Won-Gyu","Alvarez-Ponce, David"],"year":2025,"date_issued":"2025","date_published":"2025","updated_at":"2026-07-27T21:45:57Z","subjects":["Bioinformatics pipelines","Genome Annotation","Genomics","Large Language Models"],"languages":["en_US","English"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarwolf.unr.edu/handle/11714/11430","outbound_label":"Repository record","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Cushman, John C.","Yim, Won C."]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Harper, Jeffrey F.","Choi, Won-Gyu","Alvarez-Ponce, David"]},{"key":"dc:creator","label":"Author","values":["Lomas, Johnathan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-07-02T18:46:17Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-07-02T18:46:17Z"]},{"key":"dc:date.issued","label":"Date","values":["2025"]},{"key":"dc:type","label":"Dc Type","values":["Dissertation"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctorate Degree"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bioinformatics pipelines","Genome Annotation","Genomics","Large Language Models"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://scholarwolf.unr.edu/handle/11714/11430"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The annotation of protein-coding genes from a raw genome assembly involves identifying the sequence regions that are transcribed into mRNA transcripts, spliced, and ultimately translated into protein. Thus, genome annotation is a fundamental task that critically supports systems level biological inquiry by providing a reference for 'omics' analyses including, RNA sequencing, proteomics, and comparative genomics. Despite advances in genome sequencing, which enable routine chromosome-level genome assembly, complete identification of gene structures in eukaryotes remains a significant challenge. For example, existing genome annotation pipelines suffer from a lack of automation, an inability to control precision, and from poor performance in predicting alternative splicing.The goal of this work is to develop improved computational tools for genome annotation and to apply the improved methods to annotate the genome of cactus pear (Opuntia cochenillifera). To address challenges in computational efficiency and precision, an automated bioinformatics pipeline called Sylvan was developed that computes a comprehensive genome annotation from disparate evidence sources and filters spurious gene models using a semi-supervised random forest classifier. In benchmarking trials involving Arabidopsis thaliana and Oryza sativa the pipeline outperformed current standards, such as MAKER and BRAKER, in both F1 similarity and BUSCO completeness. Sylvan was used to annotate the genome of Opuntia cochenillifera, representing the first genome sequence and assembly in the genus and a foundational tool for research into crassulacean acid metabolism (CAM) and drought tolerance in plants. To improve the capacity of genome annotation tools to predict full-length, alternatively spliced transcripts ab initio, a deep learning transformer model was developed to 'translate' a DNA sequence into its text-based annotation. This generative strategy provides increased flexibility to predict hierarchical and overlapping gene structures that are not possible with one dimensional segmentation models."]},{"key":"dc:format","label":"Dc Format","values":["PDF"]},{"key":"dc:title","label":"Title","values":["Decoding DNA: scalable genome annotation software, generative AI, and application to the cactus pear genome"]}]}],"canonical_facts":{"dc:contributor.advisor":["Cushman, John C.","Yim, Won C."],"dc:contributor.committeemember":["Harper, Jeffrey F.","Choi, Won-Gyu","Alvarez-Ponce, David"],"dc:creator":["Lomas, Johnathan"],"dc:date.accessioned":["2025-07-02T18:46:17Z"],"dc:date.available":["2025-07-02T18:46:17Z"],"dc:date.issued":["2025"],"dc:description.abstract":["The annotation of protein-coding genes from a raw genome assembly involves identifying the sequence regions that are transcribed into mRNA transcripts, spliced, and ultimately translated into protein. Thus, genome annotation is a fundamental task that critically supports systems level biological inquiry by providing a reference for 'omics' analyses including, RNA sequencing, proteomics, and comparative genomics. Despite advances in genome sequencing, which enable routine chromosome-level genome assembly, complete identification of gene structures in eukaryotes remains a significant challenge. For example, existing genome annotation pipelines suffer from a lack of automation, an inability to control precision, and from poor performance in predicting alternative splicing.The goal of this work is to develop improved computational tools for genome annotation and to apply the improved methods to annotate the genome of cactus pear (Opuntia cochenillifera). To address challenges in computational efficiency and precision, an automated bioinformatics pipeline called Sylvan was developed that computes a comprehensive genome annotation from disparate evidence sources and filters spurious gene models using a semi-supervised random forest classifier. In benchmarking trials involving Arabidopsis thaliana and Oryza sativa the pipeline outperformed current standards, such as MAKER and BRAKER, in both F1 similarity and BUSCO completeness. Sylvan was used to annotate the genome of Opuntia cochenillifera, representing the first genome sequence and assembly in the genus and a foundational tool for research into crassulacean acid metabolism (CAM) and drought tolerance in plants. To improve the capacity of genome annotation tools to predict full-length, alternatively spliced transcripts ab initio, a deep learning transformer model was developed to 'translate' a DNA sequence into its text-based annotation. This generative strategy provides increased flexibility to predict hierarchical and overlapping gene structures that are not possible with one dimensional segmentation models."],"dc:format":["PDF"],"dc:identifier.uri":["https://scholarwolf.unr.edu/handle/11714/11430"],"dc:language":["English"],"dc:language.iso":["en_US"],"dc:subject":["Bioinformatics pipelines","Genome Annotation","Genomics","Large Language Models"],"dc:title":["Decoding DNA: scalable genome annotation software, generative AI, and application to the cactus pear genome"],"dc:type":["Dissertation"],"thesis:degree_level":["Doctorate Degree"]},"updated_at":"2026-07-27T21:45:57Z"}