Abstract
dc:description.abstract<p>Data clustering plays an important role in effective analysis of gene expression. Although DNA microarray technology facilitates expression monitoring, several challenges arise when dealing with gene expression datasets. Some of these challenges are the enormous number of genes, the dimensionality of the data, and the change of data over time. The genetic groups which are biologically interlinked can be identified through clustering. This project aims to clarify the steps to apply clustering analysis of genes involved in a published dataset. The methodology for this project includes the selection of the dataset representation, the selection of gene datasets, Similarity Matrix Selection, the selection of clustering algorithm, and analysis tool. R language with the focus of Kmeans, fpc, hclust, and heatmap3 packages in R is used in this project as an analysis tool. Different clustering algorithms are used on Spellman dataset to illustrate how genes are grouped together in clusters which help to understand our genetic behaviors.</p>
Degree
thesis:*- Name thesis:degree_name
- Master of Science in Computer Science
- Level thesis:degree_level
- Project
- Discipline thesis:degree_discipline
- School of Computer Science and Engineering
- Year dc:date.available
- 2015
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Abualhamayl, Abdullah Jameel, Mr.
- Contributors dc:contributor
-
- Haiyan Qiao
Subjects
dc:subject × 9Identifiers
dc:identifier.*- Repository record dc:identifier
- https://scholarworks.lib.csusb.edu/etd/259
- OAI identifier oai:identifier
- oai:scholarworks.lib.csusb.edu:etd-1293