Abstract
dc:description.abstractKnowledge discovery, in databases, also known as data mining, is aimed to find significant information from a set of data. The knowledge to be mined from the dataset may refer to patterns, association rules, classification and clustering rules, and so forth. In this dissertation, we present a neural network approach to finding knowledge in biological databases. Specifically, we propose new methods to process biological sequences in two case studies: the classification of protein sequences and the prediction of E. Coli promoters in DNA sequences. Our proposed methods, based oil neural network architectures combine techniques ranging from Bayesian inference, coding theory, feature selection, dimensionality reduction, to dynamic programming and machine learning algorithms. Empirical studies show that the proposed methods outperform previously published methods and have excellent performance on the latest dataset. We have implemented the proposed algorithms into an infrastructure, called Genome Mining, developed for biosequence classification and recognition.
Degree
thesis:*- Name thesis:degree_name
- Doctor of Philosophy in Computing Sciences - (Ph.D.)
- Discipline thesis:degree_discipline
- Computer and Information Science
- Year
- 2000
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Ma, Qicheng
- Contributors dc:contributor
-
- Jason T. L. Wang
- James A. McHugh
- Frank Y. Shih
Subjects
dc:subject × 6Identifiers
dc:identifier.*- Repository record dc:identifier
- https://digitalcommons.njit.edu/dissertations/423
- OAI identifier oai:identifier
- oai:digitalcommons.njit.edu:dissertations-1478