Back to search

George Mason University

Classification of Thermophilic and Mesophilic Proteins Using N-Grams

Abstract

The project is focused on machine learning classification of thermophilic and mesophilic proteins using N-gram based representation of protein sequences. Two datasets containing proteins from both classes were used for the analysis. Alphabet reduction was performed on all datasets, and n-gram frequencies were calculated for each sequence using the reduced alphabet. Data normalization was done by calculating n-gram likelihoods. Four different machine learning algorithms (Naïve Bayes, Support Vector Machines, Decision Trees and Random Forests) were used for the protein classification. Accuracies of 100.0% were achieved using SVM, 99.3% using Random Forests, 90.3% using Naïve Bayes and 99.6 using Decision Trees.

Author and committee

dc:creator, dc:contributor.*
Author
  • Elattar, Marwy

Subjects

dc:subject × 4

Identifiers

dc:identifier.*
Identifier
hdl:1920/10794
OAI identifier oai:identifier
oai:MARS:1920/10794

Chain of custody

source
Harvested from
George Mason University
Base URL
mars.gmu.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Elattar, Marwy. Classification of Thermophilic and Mesophilic Proteins Using N-Grams.