Back to results

York University

Using signal processing, evolutionary computation, and machine learning to identify transposable elements in genomes

Abstract

dc:description.abstract

About half of the human genome consists of transposable elements (TE's), sequences that have many copies of themselves distributed throughout the genome. All genomes, from bacterial to human, contain TE's. TE's affect genome function by either creating proteins directly or affecting genome regulation. They serve as molecular fossils, giving clues to the evolutionary history of the organism. TE's are often challenging to identify because they are fragmentary or heavily mutated. In this thesis, novel features for the detection and study of TE's are developed. These features are of two types. The first type are statistical features based on the Fourier transform used to assess reading frame use. These features measure how different the reading frame use is from that of a random sequence, which reading frames the sequence is using, and the proportion of use of the active reading frames. The second type of feature, called side effect machine (SEM) features, are generated by finite state machines augmented with counters that track the number of times the state is visited. These counters then become features of the sequence. The number of possible SEM features is super-exponential in the number of states. New methods for selecting useful feature subsets that incorporate a genetic algorithm and a novel clustering method are introduced. The features produced reveal structural characteristics of the sequences of potential interest to biologists. A detailed analysis of the genetic algorithm, its fitness functions, and its fitness landscapes is performed. The features are used, together with features used in existing exon finding algorithms, to build classifiers that distinguish TE's from other genomic sequences in humans, fruit flies, and ciliates. The classifiers achieve high accuracy (> 85%) on a variety of TE classification problems. The classifiers are used to scan large genomes for TE's. In addition, the features are used to describe the TE's in the newly sequenced ciliate, Tetrahymena thermophile to provide information for biologists useful to them in forming hypotheses to test experimentally concerning the role of these TE's and the mechanisms that govern them.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Ashlock, Wendy Cole
Advisor dc:contributor.advisor
  • Datta, Suprakash

Rights

dc:rights
Statement dc:rights
  • Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests.

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/10315/31432
OAI identifier oai:identifier
oai:yorkspace.library.yorku.ca:10315/31432

Chain of custody

source
Harvested from
York University
Base URL
yorkspace.library.yorku.ca/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Ashlock, Wendy Cole. Using signal processing, evolutionary computation, and machine learning to identify transposable elements in genomes. 2016. http://hdl.handle.net/10315/31432