{"id":{"repo_id":"mississippi","oai_identifier":"oai:egrove.olemiss.edu:etd-1444"},"canonical_url":"https://search.dev.ndltd.org/etd/mississippi/oai:egrove.olemiss.edu:etd-1444","repository":{"repo_id":"mississippi","name":"University of Mississippi","base_url":"https://egrove.olemiss.edu/do/oai/"},"display":{"title":"Large Margin Random Forests On Mixed Type Data","abstract":"Incorporating various sources of biological information is important for biological discovery. For example, genes have a multi-view representation. They can be represented by features such as sequence length and physical-chemical properties. They can also be represented by pairwise similarities, gene expression levels, and phylogenetics position. Hence, the types vary from numerical features to categorical features. An efficient way of learning from observations with a multi-view representation of mixed type of data is thus important. We propose a large margin random forests classification approach based on random forests proximity. Random forests accommodate mixed data types naturally. Large margin classifiers are obtained from the random forests proximity kernel or its derivative kernels. We test the approach on four biological datasets. The performance is promising compared with other state of the art methods including support vector machines (SVMs) and Random Forests classifiers. It demonstrates high potential in the discovery of functional roles of genes and proteins. We also examine the effects of mixed type of data on the algorithms used.","abstract_html":"Incorporating various sources of biological information is important for biological discovery. For example, genes have a multi-view representation. They can be represented by features such as sequence length and physical-chemical properties. They can also be represented by pairwise similarities, gene expression levels, and phylogenetics position. Hence, the types vary from numerical features to categorical features. An efficient way of learning from observations with a multi-view representation of mixed type of data is thus important. We propose a large margin random forests classification approach based on random forests proximity. Random forests accommodate mixed data types naturally. Large margin classifiers are obtained from the random forests proximity kernel or its derivative kernels. We test the approach on four biological datasets. The performance is promising compared with other state of the art methods including support vector machines (SVMs) and Random Forests classifiers. It demonstrates high potential in the discovery of functional roles of genes and proteins. We also examine the effects of mixed type of data on the algorithms used.","abstract_has_math":false,"creators":["Liu, Sheng"],"institution":null,"degree_name":"M.S. in Engineering Science","degree_level":"Thesis","degree_discipline":"Computer and Information Science","degree_department":null,"school":null,"contributors":["Yixin Chen","Conrad Cunningham"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-01-01T08:00:00Z","date_published":"2011-01-01T08:00:00Z","updated_at":"2026-07-24T03:05:36Z","subjects":["Computer Sciences"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://egrove.olemiss.edu/etd/445","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Yixin Chen","Conrad Cunningham"]},{"key":"dc:creator","label":"Author","values":["Liu, Sheng"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2019-06-20T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer and Information Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S. in Engineering Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Sciences"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://egrove.olemiss.edu/etd/445"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Incorporating various sources of biological information is important for biological discovery. For example, genes have a multi-view representation. They can be represented by features such as sequence length and physical-chemical properties. They can also be represented by pairwise similarities, gene expression levels, and phylogenetics position. Hence, the types vary from numerical features to categorical features. An efficient way of learning from observations with a multi-view representation of mixed type of data is thus important. We propose a large margin random forests classification approach based on random forests proximity. Random forests accommodate mixed data types naturally. Large margin classifiers are obtained from the random forests proximity kernel or its derivative kernels. We test the approach on four biological datasets. The performance is promising compared with other state of the art methods including support vector machines (SVMs) and Random Forests classifiers. It demonstrates high potential in the discovery of functional roles of genes and proteins. We also examine the effects of mixed type of data on the algorithms used."]},{"key":"dc:title","label":"Title","values":["Large Margin Random Forests On Mixed Type Data"]}]}],"canonical_facts":{"dc:contributor":["Yixin Chen","Conrad Cunningham"],"dc:creator":["Liu, Sheng"],"dc:date.available":["2019-06-20T07:00:00Z"],"dc:description.abstract":["Incorporating various sources of biological information is important for biological discovery. For example, genes have a multi-view representation. They can be represented by features such as sequence length and physical-chemical properties. They can also be represented by pairwise similarities, gene expression levels, and phylogenetics position. Hence, the types vary from numerical features to categorical features. An efficient way of learning from observations with a multi-view representation of mixed type of data is thus important. We propose a large margin random forests classification approach based on random forests proximity. Random forests accommodate mixed data types naturally. Large margin classifiers are obtained from the random forests proximity kernel or its derivative kernels. We test the approach on four biological datasets. The performance is promising compared with other state of the art methods including support vector machines (SVMs) and Random Forests classifiers. It demonstrates high potential in the discovery of functional roles of genes and proteins. We also examine the effects of mixed type of data on the algorithms used."],"dc:identifier":["https://egrove.olemiss.edu/etd/445"],"dc:subject":["Computer Sciences"],"dc:title":["Large Margin Random Forests On Mixed Type Data"],"thesis:degree_discipline":["Computer and Information Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S. in Engineering Science"]},"updated_at":"2026-07-24T03:05:36Z"}