Back to results

University of Nevada - Reno

Decoding DNA: scalable genome annotation software, generative AI, and application to the cactus pear genome

Abstract

dc:description.abstract

The annotation of protein-coding genes from a raw genome assembly involves identifying the sequence regions that are transcribed into mRNA transcripts, spliced, and ultimately translated into protein. Thus, genome annotation is a fundamental task that critically supports systems level biological inquiry by providing a reference for 'omics' analyses including, RNA sequencing, proteomics, and comparative genomics. Despite advances in genome sequencing, which enable routine chromosome-level genome assembly, complete identification of gene structures in eukaryotes remains a significant challenge. For example, existing genome annotation pipelines suffer from a lack of automation, an inability to control precision, and from poor performance in predicting alternative splicing.The goal of this work is to develop improved computational tools for genome annotation and to apply the improved methods to annotate the genome of cactus pear (Opuntia cochenillifera). To address challenges in computational efficiency and precision, an automated bioinformatics pipeline called Sylvan was developed that computes a comprehensive genome annotation from disparate evidence sources and filters spurious gene models using a semi-supervised random forest classifier. In benchmarking trials involving Arabidopsis thaliana and Oryza sativa the pipeline outperformed current standards, such as MAKER and BRAKER, in both F1 similarity and BUSCO completeness. Sylvan was used to annotate the genome of Opuntia cochenillifera, representing the first genome sequence and assembly in the genus and a foundational tool for research into crassulacean acid metabolism (CAM) and drought tolerance in plants. To improve the capacity of genome annotation tools to predict full-length, alternatively spliced transcripts ab initio, a deep learning transformer model was developed to 'translate' a DNA sequence into its text-based annotation. This generative strategy provides increased flexibility to predict hierarchical and overlapping gene structures that are not possible with one dimensional segmentation models.

Degree

thesis:*
Level thesis:degree_level
Doctorate Degree
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Lomas, Johnathan
Advisors dc:contributor.advisor
  • Cushman, John C.
  • Yim, Won C.
Committee members dc:contributor.committeemember
  • Harper, Jeffrey F.
  • Choi, Won-Gyu
  • Alvarez-Ponce, David

Subjects

dc:subject × 4

Rights

Language dc:language.iso
en_US, English

Identifiers

dc:identifier.*
Repository record dc:identifier.uri
https://scholarwolf.unr.edu/handle/11714/11430
OAI identifier oai:identifier
oai:scholarwolf.unr.edu:11714/11430

Chain of custody

source
Harvested from
University of Nevada - Reno
Base URL
scholarwolf.unr.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Lomas, Johnathan. Decoding DNA: scalable genome annotation software, generative AI, and application to the cactus pear genome. Doctorate Degree thesis, 2025. https://scholarwolf.unr.edu/handle/11714/11430