{"id":{"repo_id":"ohiolink","oai_identifier":"oai:etd.ohiolink.edu:case1365089538"},"canonical_url":"https://search.dev.ndltd.org/etd/ohiolink/oai:etd.ohiolink.edu:case1365089538","repository":{"repo_id":"ohiolink","name":"OhioLINK","base_url":"https://etd.ohiolink.edu/acprod/odb_etd/ws/oai/oai"},"display":{"title":"Motif Mining On Structured And Semi-structured Biological Data","abstract":"Motif discovery in biological data is very important to biologists becausethey are conjectured to have biological significance. Despite the considerateeffort put in this topic, it remains a challenging and difficult problem.There are several advances in biology such that the data is structured (graphs)and semi-structured (sequences).A challenge of motif mining in sequences is the existence of variationsincluding substitutions and permutations. Taking into account the existenceof these two kinds of variations, we propose a novel sequential motif modeland devise an algorithm to discover those motifs. A reachability propertyis identified to prune the search space. We also demonstrate that the discoveredmotifs have biological significance. Motivated by another biologicalphenomenal called compensational mutation in biological sequences, we proposeanother model for motif discovery in biological sequences. Consideringcompensational mutation adds more degree of freedom in the search space.A novel algorithm is proposed to solve the problem. It is shown that theproposed algorithm can discover effective motifs in a timely manner.Besides sequential data, a huge amount of biological data can be naturallyrepresented as graphs, e.g., protein interaction networks and gene regulatorynetworks. Most of the existing graph mining research focuses on miningunweighted graphs. However, weighted graphs are actually more common. Toaddress the problem of motif discovery in weighted graphs, a weighted subgraphmotif model is proposed to capture the importance of a subgraph motifin a single large weighted graph. We study two related problems about subgraphmotif discovery in a large weighted graph: (1) discovering all motifs withrespect to a given minimum weight threshold and (2) finding top k motifs withthe largest weights. We identify a property called 1-extension property so thata bounded search can be achieved. Last but not least, real and synthetic datasets are used to show the effectiveness and efficiency of the proposed modeland algorithm.","abstract_html":"Motif discovery in biological data is very important to biologists becausethey are conjectured to have biological significance. Despite the considerateeffort put in this topic, it remains a challenging and difficult problem.There are several advances in biology such that the data is structured (graphs)and semi-structured (sequences).A challenge of motif mining in sequences is the existence of variationsincluding substitutions and permutations. Taking into account the existenceof these two kinds of variations, we propose a novel sequential motif modeland devise an algorithm to discover those motifs. A reachability propertyis identified to prune the search space. We also demonstrate that the discoveredmotifs have biological significance. Motivated by another biologicalphenomenal called compensational mutation in biological sequences, we proposeanother model for motif discovery in biological sequences. Consideringcompensational mutation adds more degree of freedom in the search space.A novel algorithm is proposed to solve the problem. It is shown that theproposed algorithm can discover effective motifs in a timely manner.Besides sequential data, a huge amount of biological data can be naturallyrepresented as graphs, e.g., protein interaction networks and gene regulatorynetworks. Most of the existing graph mining research focuses on miningunweighted graphs. However, weighted graphs are actually more common. Toaddress the problem of motif discovery in weighted graphs, a weighted subgraphmotif model is proposed to capture the importance of a subgraph motifin a single large weighted graph. We study two related problems about subgraphmotif discovery in a large weighted graph: (1) discovering all motifs withrespect to a given minimum weight threshold and (2) finding top k motifs withthe largest weights. We identify a property called 1-extension property so thata bounded search can be achieved. Last but not least, real and synthetic datasets are used to show the effectiveness and efficiency of the proposed modeland algorithm.","abstract_has_math":false,"creators":["Su, Wei"],"institution":"Case Western Reserve University School of Graduate Studies","degree_name":"Doctor of Philosophy","degree_level":"doctoral","degree_discipline":"EECS - Computer and Information Sciences","degree_department":null,"school":null,"contributors":["Koyuturk, Mehmet","Yang, Jiong"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2013,"date_issued":"2013-08-19","date_published":"2013-08-19","updated_at":"2026-07-24T03:37:16Z","subjects":["Computer Science"],"languages":["English"],"rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://rave.ohiolink.edu/etdc/view?acc_num=case1365089538","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Koyuturk, Mehmet","Yang, Jiong"]},{"key":"dc:creator","label":"Author","values":["Su, Wei"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2013-08-19"]},{"key":"dc:publisher","label":"Institution","values":["Case Western Reserve University School of Graduate Studies / OhioLINK"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["EECS - Computer and Information Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Case Western Reserve University School of Graduate Studies"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer Science"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["English"]},{"key":"dc:rights","label":"Dc Rights","values":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://rave.ohiolink.edu/etdc/view?acc_num=case1365089538"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Motif discovery in biological data is very important to biologists becausethey are conjectured to have biological significance. Despite the considerateeffort put in this topic, it remains a challenging and difficult problem.There are several advances in biology such that the data is structured (graphs)and semi-structured (sequences).A challenge of motif mining in sequences is the existence of variationsincluding substitutions and permutations. Taking into account the existenceof these two kinds of variations, we propose a novel sequential motif modeland devise an algorithm to discover those motifs. A reachability propertyis identified to prune the search space. We also demonstrate that the discoveredmotifs have biological significance. Motivated by another biologicalphenomenal called compensational mutation in biological sequences, we proposeanother model for motif discovery in biological sequences. Consideringcompensational mutation adds more degree of freedom in the search space.A novel algorithm is proposed to solve the problem. It is shown that theproposed algorithm can discover effective motifs in a timely manner.Besides sequential data, a huge amount of biological data can be naturallyrepresented as graphs, e.g., protein interaction networks and gene regulatorynetworks. Most of the existing graph mining research focuses on miningunweighted graphs. However, weighted graphs are actually more common. Toaddress the problem of motif discovery in weighted graphs, a weighted subgraphmotif model is proposed to capture the importance of a subgraph motifin a single large weighted graph. We study two related problems about subgraphmotif discovery in a large weighted graph: (1) discovering all motifs withrespect to a given minimum weight threshold and (2) finding top k motifs withthe largest weights. We identify a property called 1-extension property so thata bounded search can be achieved. Last but not least, real and synthetic datasets are used to show the effectiveness and efficiency of the proposed modeland algorithm."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf","p.135","771.9 KB"]},{"key":"dc:title","label":"Title","values":["Motif Mining On Structured And Semi-structured Biological Data"]}]}],"canonical_facts":{"dc:contributor":["Koyuturk, Mehmet","Yang, Jiong"],"dc:creator":["Su, Wei"],"dc:date":["2013-08-19"],"dc:description":["Motif discovery in biological data is very important to biologists becausethey are conjectured to have biological significance. Despite the considerateeffort put in this topic, it remains a challenging and difficult problem.There are several advances in biology such that the data is structured (graphs)and semi-structured (sequences).A challenge of motif mining in sequences is the existence of variationsincluding substitutions and permutations. Taking into account the existenceof these two kinds of variations, we propose a novel sequential motif modeland devise an algorithm to discover those motifs. A reachability propertyis identified to prune the search space. We also demonstrate that the discoveredmotifs have biological significance. Motivated by another biologicalphenomenal called compensational mutation in biological sequences, we proposeanother model for motif discovery in biological sequences. Consideringcompensational mutation adds more degree of freedom in the search space.A novel algorithm is proposed to solve the problem. It is shown that theproposed algorithm can discover effective motifs in a timely manner.Besides sequential data, a huge amount of biological data can be naturallyrepresented as graphs, e.g., protein interaction networks and gene regulatorynetworks. Most of the existing graph mining research focuses on miningunweighted graphs. However, weighted graphs are actually more common. Toaddress the problem of motif discovery in weighted graphs, a weighted subgraphmotif model is proposed to capture the importance of a subgraph motifin a single large weighted graph. We study two related problems about subgraphmotif discovery in a large weighted graph: (1) discovering all motifs withrespect to a given minimum weight threshold and (2) finding top k motifs withthe largest weights. We identify a property called 1-extension property so thata bounded search can be achieved. Last but not least, real and synthetic datasets are used to show the effectiveness and efficiency of the proposed modeland algorithm."],"dc:format":["application/pdf","p.135","771.9 KB"],"dc:identifier":["http://rave.ohiolink.edu/etdc/view?acc_num=case1365089538"],"dc:language":["English"],"dc:publisher":["Case Western Reserve University School of Graduate Studies / OhioLINK"],"dc:rights":["unrestricted","This thesis or dissertation is protected by copyright: all rights reserved. It may not be copied or redistributed beyond the terms of applicable copyright laws."],"dc:subject":["Computer Science"],"dc:title":["Motif Mining On Structured And Semi-structured Biological Data"],"dc:type":["Electronic Thesis or Dissertation"],"thesis:degree_discipline":["EECS - Computer and Information Sciences"],"thesis:degree_level":["doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["Case Western Reserve University School of Graduate Studies"]},"updated_at":"2026-07-24T03:37:16Z"}