{"id":{"repo_id":"umkc","oai_identifier":"oai:mospace.umsystem.edu:10355/60055"},"canonical_url":"https://search.dev.ndltd.org/etd/umkc/oai:mospace.umsystem.edu:10355/60055","repository":{"repo_id":"umkc","name":"University of Missouri - Kansas City","base_url":"https://mospace.umsystem.edu/oai/request"},"display":{"title":"PPDQ-BG: Parallel Partition and Distributed Query Processing for Big Graphs","abstract":"In recent years, there has been an explosive growth of the linked data of a global information space that often requires expensive computations to perform big graph analysis and query processing. Graph data represent irregular and unstructured relationships that usually result in a lack of locality so that it is often difficult to extract relevant information from big graphs. Although there have been advances in graph processing in centralized, as well as distributed environments, there was a lack of an efficient method of handling large scale graphs for the task of finding relevant relations from big graphs or partitioning a big graph into several meaningful inter-connected graph partitions. In this thesis, we propose a scalable framework, “Parallel Partition and Distributed Query Processing for Big Graphs” (PPDQ-BG) that aims to achieve a parallel partition of large scale RDF graph and distributed query processing on the partitioned data. In this thesis, we propose a PPDQ-BG framework to make the previous centralized approach, “A Big Graph Analytics Framework for Knowledge Discovery”, a distributed model by proposing new partitioning algorithms. The proposed framework also has parallel computation of relevant information by determining neighborhood relationships among predicates of large RDF graphs in a parallel manner by computing the similarity of predicates in large scale graphs. It has the design of parallel algorithms for partitioning of graphs in a distributed dataflow framework by proposing new clustering algorithms, Similarity based Fuzzy C-Means Partitioning and Hierarchical Predicate Based Clustering. The framework also has implementation of an interactive tool using Neo4j graph databases for executing a distributed query process from partitioned graphs, experimental evaluations including correlation coefficient matrix evaluations to validate the proposed framework and comparison of proposed partitioning methods with existing partitioning algorithms using multiple datasets including medical ontology datasets, DBPedia, YAGO, and Bio2RDF datasets, experimental results of distributed query processing for efficient data retrieval for various complex queries against large scale datasets including DBPedia, YAGO, and Bio2RDF datasets.","abstract_html":"In recent years, there has been an explosive growth of the linked data of a global information space that often requires expensive computations to perform big graph analysis and query processing. Graph data represent irregular and unstructured relationships that usually result in a lack of locality so that it is often difficult to extract relevant information from big graphs. Although there have been advances in graph processing in centralized, as well as distributed environments, there was a lack of an efficient method of handling large scale graphs for the task of finding relevant relations from big graphs or partitioning a big graph into several meaningful inter-connected graph partitions. In this thesis, we propose a scalable framework, “Parallel Partition and Distributed Query Processing for Big Graphs” (PPDQ-BG) that aims to achieve a parallel partition of large scale RDF graph and distributed query processing on the partitioned data. In this thesis, we propose a PPDQ-BG framework to make the previous centralized approach, “A Big Graph Analytics Framework for Knowledge Discovery”, a distributed model by proposing new partitioning algorithms. The proposed framework also has parallel computation of relevant information by determining neighborhood relationships among predicates of large RDF graphs in a parallel manner by computing the similarity of predicates in large scale graphs. It has the design of parallel algorithms for partitioning of graphs in a distributed dataflow framework by proposing new clustering algorithms, Similarity based Fuzzy C-Means Partitioning and Hierarchical Predicate Based Clustering. The framework also has implementation of an interactive tool using Neo4j graph databases for executing a distributed query process from partitioned graphs, experimental evaluations including correlation coefficient matrix evaluations to validate the proposed framework and comparison of proposed partitioning methods with existing partitioning algorithms using multiple datasets including medical ontology datasets, DBPedia, YAGO, and Bio2RDF datasets, experimental results of distributed query processing for efficient data retrieval for various complex queries against large scale datasets including DBPedia, YAGO, and Bio2RDF datasets.","abstract_has_math":false,"creators":["Kandula, Lema"],"institution":"University of Missouri--Kansas City","degree_name":"M.S.","degree_level":"Masters","degree_discipline":"Computer Science (UMKC)","degree_department":null,"school":null,"contributors":[],"advisors":["Lee, Yugyung, 1960-"],"committee_chairs":[],"committee_members":[],"year":2016,"date_issued":"2016","date_published":"2016","updated_at":"2026-07-24T05:18:34Z","subjects":[],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10355/60055","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lee, Yugyung, 1960-"]},{"key":"dc:creator","label":"Author","values":["Kandula, Lema"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2017-04-18T17:54:00Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2017-04-18T17:54:00Z"]},{"key":"dc:date.issued","label":"Date","values":["2016"]},{"key":"dc:publisher","label":"Institution","values":["University of Missouri--Kansas City"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science (UMKC)"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Missouri--Kansas City"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10355/60055"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Title from PDF of title page viewed April 26, 2017","Thesis advisor: Yugyung Lee","Vita","Includes bibliographical references (pages 79-82)","Thesis (M.S.)--School of Computing and Engineering. University of Missouri--Kansas City, 2016"]},{"key":"dc:description.abstract","label":"Abstract","values":["In recent years, there has been an explosive growth of the linked data of a global information space that often requires expensive computations to perform big graph analysis and query processing. Graph data represent irregular and unstructured relationships that usually result in a lack of locality so that it is often difficult to extract relevant information from big graphs. Although there have been advances in graph processing in centralized, as well as distributed environments, there was a lack of an efficient method of handling large scale graphs for the task of finding relevant relations from big graphs or partitioning a big graph into several meaningful inter-connected graph partitions. In this thesis, we propose a scalable framework, “Parallel Partition and Distributed Query Processing for Big Graphs” (PPDQ-BG) that aims to achieve a parallel partition of large scale RDF graph and distributed query processing on the partitioned data. In this thesis, we propose a PPDQ-BG framework to make the previous centralized approach, “A Big Graph Analytics Framework for Knowledge Discovery”, a distributed model by proposing new partitioning algorithms. The proposed framework also has parallel computation of relevant information by determining neighborhood relationships among predicates of large RDF graphs in a parallel manner by computing the similarity of predicates in large scale graphs. It has the design of parallel algorithms for partitioning of graphs in a distributed dataflow framework by proposing new clustering algorithms, Similarity based Fuzzy C-Means Partitioning and Hierarchical Predicate Based Clustering. The framework also has implementation of an interactive tool using Neo4j graph databases for executing a distributed query process from partitioned graphs, experimental evaluations including correlation coefficient matrix evaluations to validate the proposed framework and comparison of proposed partitioning methods with existing partitioning algorithms using multiple datasets including medical ontology datasets, DBPedia, YAGO, and Bio2RDF datasets, experimental results of distributed query processing for efficient data retrieval for various complex queries against large scale datasets including DBPedia, YAGO, and Bio2RDF datasets."]},{"key":"dc:title","label":"Title","values":["PPDQ-BG: Parallel Partition and Distributed Query Processing for Big Graphs"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lee, Yugyung, 1960-"],"dc:creator":["Kandula, Lema"],"dc:date.accessioned":["2017-04-18T17:54:00Z"],"dc:date.available":["2017-04-18T17:54:00Z"],"dc:date.issued":["2016"],"dc:description":["Title from PDF of title page viewed April 26, 2017","Thesis advisor: Yugyung Lee","Vita","Includes bibliographical references (pages 79-82)","Thesis (M.S.)--School of Computing and Engineering. University of Missouri--Kansas City, 2016"],"dc:description.abstract":["In recent years, there has been an explosive growth of the linked data of a global information space that often requires expensive computations to perform big graph analysis and query processing. Graph data represent irregular and unstructured relationships that usually result in a lack of locality so that it is often difficult to extract relevant information from big graphs. Although there have been advances in graph processing in centralized, as well as distributed environments, there was a lack of an efficient method of handling large scale graphs for the task of finding relevant relations from big graphs or partitioning a big graph into several meaningful inter-connected graph partitions. In this thesis, we propose a scalable framework, “Parallel Partition and Distributed Query Processing for Big Graphs” (PPDQ-BG) that aims to achieve a parallel partition of large scale RDF graph and distributed query processing on the partitioned data. In this thesis, we propose a PPDQ-BG framework to make the previous centralized approach, “A Big Graph Analytics Framework for Knowledge Discovery”, a distributed model by proposing new partitioning algorithms. The proposed framework also has parallel computation of relevant information by determining neighborhood relationships among predicates of large RDF graphs in a parallel manner by computing the similarity of predicates in large scale graphs. It has the design of parallel algorithms for partitioning of graphs in a distributed dataflow framework by proposing new clustering algorithms, Similarity based Fuzzy C-Means Partitioning and Hierarchical Predicate Based Clustering. The framework also has implementation of an interactive tool using Neo4j graph databases for executing a distributed query process from partitioned graphs, experimental evaluations including correlation coefficient matrix evaluations to validate the proposed framework and comparison of proposed partitioning methods with existing partitioning algorithms using multiple datasets including medical ontology datasets, DBPedia, YAGO, and Bio2RDF datasets, experimental results of distributed query processing for efficient data retrieval for various complex queries against large scale datasets including DBPedia, YAGO, and Bio2RDF datasets."],"dc:identifier.uri":["https://hdl.handle.net/10355/60055"],"dc:language.iso":["en_US"],"dc:publisher":["University of Missouri--Kansas City"],"dc:title":["PPDQ-BG: Parallel Partition and Distributed Query Processing for Big Graphs"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science (UMKC)"],"thesis:degree_level":["Masters"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Missouri--Kansas City"]},"updated_at":"2026-07-24T05:18:34Z"}