{"id":{"repo_id":"houston","oai_identifier":"oai:uh-ir.tdl.org:10657/17736"},"canonical_url":"https://search.dev.ndltd.org/etd/houston/oai:uh-ir.tdl.org:10657/17736","repository":{"repo_id":"houston","name":"University of Houston","base_url":"https://uh-ir.tdl.org/server/oai/request"},"display":{"title":"Graph Analytics with Data Science Languages","abstract":"In Big Data analytics, data exploded in the three Vs: Volume, Velocity, and Variety. The three Vs brought new challenges to data analysis systems, which require new approaches, tools, and algorithms to analyze data. The most complex exploration mechanism in Big Data Analytics is graphs, which are flexible to represent any set of interconnected objects. Graph analytics is particularly challenging mainly due to large graph sizes and the structure of graphs. In this dissertation, we work on analyzing large graphs that cannot fit in main memory. First, we extract a general computation pattern for several graph properties that can be solved with iterative algorithms. Then, we start with database management systems (DBMSs) to compute those graph properties since a lot of data stored in DBMSs can be analyzed as graphs. We proposed algorithms and optimized query solutions with a focus on graph partitions. Experimental evaluations demonstrate that our solutions can work on large graphs with good speed up and reasonable performance. After conducting a comprehensive survey about Data Science Languages, we continued our work on analyzing large graphs with Data Science Language, Python. We developed a lightweight C++ function that can be used for several graph algorithms. The function is easily called in Python. Comparing our function with other state-of-the-art graph libraries shows its good performance. Finally, we study the crucial graph algorithm, transitive closure. We propose a disk-based transitive closure solution that operates on the bit-matrix. The solution is within the Python ecosystem while adhering to the principles of database systems. Our experimental study shows the superiority of our solutions over existing popular analysis systems, suggesting potential advancements in bridging high-performance computing and Python.","abstract_html":"In Big Data analytics, data exploded in the three Vs: Volume, Velocity, and Variety. The three Vs brought new challenges to data analysis systems, which require new approaches, tools, and algorithms to analyze data. The most complex exploration mechanism in Big Data Analytics is graphs, which are flexible to represent any set of interconnected objects. Graph analytics is particularly challenging mainly due to large graph sizes and the structure of graphs. In this dissertation, we work on analyzing large graphs that cannot fit in main memory. First, we extract a general computation pattern for several graph properties that can be solved with iterative algorithms. Then, we start with database management systems (DBMSs) to compute those graph properties since a lot of data stored in DBMSs can be analyzed as graphs. We proposed algorithms and optimized query solutions with a focus on graph partitions. Experimental evaluations demonstrate that our solutions can work on large graphs with good speed up and reasonable performance. After conducting a comprehensive survey about Data Science Languages, we continued our work on analyzing large graphs with Data Science Language, Python. We developed a lightweight C++ function that can be used for several graph algorithms. The function is easily called in Python. Comparing our function with other state-of-the-art graph libraries shows its good performance. Finally, we study the crucial graph algorithm, transitive closure. We propose a disk-based transitive closure solution that operates on the bit-matrix. The solution is within the Python ecosystem while adhering to the principles of database systems. Our experimental study shows the superiority of our solutions over existing popular analysis systems, suggesting potential advancements in bridging high-performance computing and Python.","abstract_has_math":false,"creators":["Zhou, Xiantian"],"institution":"University of Houston","degree_name":"Doctor of Philosophy","degree_level":"Doctoral","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Ordonez, Carlos"],"committee_chairs":[],"committee_members":["Azencott, Robert","Subhlok, Jaspal","Huang, Stephen"],"year":2024,"date_issued":"2024-03-28","date_published":"2024-03-28","updated_at":"2026-07-24T02:32:54Z","subjects":["Graph analytics","Big data","Graph metrics","Graph patterns"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10657/17736","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Ordonez, Carlos"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Azencott, Robert","Subhlok, Jaspal","Huang, Stephen"]},{"key":"dc:creator","label":"Author","values":["Zhou, Xiantian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-07-26T22:24:29Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-07-26T22:24:29Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-03-28"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Houston"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Graph analytics","Big data","Graph metrics","Graph patterns"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10657/17736"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In Big Data analytics, data exploded in the three Vs: Volume, Velocity, and Variety. The three Vs brought new challenges to data analysis systems, which require new approaches, tools, and algorithms to analyze data. The most complex exploration mechanism in Big Data Analytics is graphs, which are flexible to represent any set of interconnected objects. Graph analytics is particularly challenging mainly due to large graph sizes and the structure of graphs. In this dissertation, we work on analyzing large graphs that cannot fit in main memory. First, we extract a general computation pattern for several graph properties that can be solved with iterative algorithms. Then, we start with database management systems (DBMSs) to compute those graph properties since a lot of data stored in DBMSs can be analyzed as graphs. We proposed algorithms and optimized query solutions with a focus on graph partitions. Experimental evaluations demonstrate that our solutions can work on large graphs with good speed up and reasonable performance. After conducting a comprehensive survey about Data Science Languages, we continued our work on analyzing large graphs with Data Science Language, Python. We developed a lightweight C++ function that can be used for several graph algorithms. The function is easily called in Python. Comparing our function with other state-of-the-art graph libraries shows its good performance. Finally, we study the crucial graph algorithm, transitive closure. We propose a disk-based transitive closure solution that operates on the bit-matrix. The solution is within the Python ecosystem while adhering to the principles of database systems. Our experimental study shows the superiority of our solutions over existing popular analysis systems, suggesting potential advancements in bridging high-performance computing and Python."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Graph Analytics with Data Science Languages"]}]}],"canonical_facts":{"dc:contributor.advisor":["Ordonez, Carlos"],"dc:contributor.committeemember":["Azencott, Robert","Subhlok, Jaspal","Huang, Stephen"],"dc:creator":["Zhou, Xiantian"],"dc:date.accessioned":["2024-07-26T22:24:29Z"],"dc:date.available":["2024-07-26T22:24:29Z"],"dc:date.issued":["2024-03-28"],"dc:description.abstract":["In Big Data analytics, data exploded in the three Vs: Volume, Velocity, and Variety. The three Vs brought new challenges to data analysis systems, which require new approaches, tools, and algorithms to analyze data. The most complex exploration mechanism in Big Data Analytics is graphs, which are flexible to represent any set of interconnected objects. Graph analytics is particularly challenging mainly due to large graph sizes and the structure of graphs. In this dissertation, we work on analyzing large graphs that cannot fit in main memory. First, we extract a general computation pattern for several graph properties that can be solved with iterative algorithms. Then, we start with database management systems (DBMSs) to compute those graph properties since a lot of data stored in DBMSs can be analyzed as graphs. We proposed algorithms and optimized query solutions with a focus on graph partitions. Experimental evaluations demonstrate that our solutions can work on large graphs with good speed up and reasonable performance. After conducting a comprehensive survey about Data Science Languages, we continued our work on analyzing large graphs with Data Science Language, Python. We developed a lightweight C++ function that can be used for several graph algorithms. The function is easily called in Python. Comparing our function with other state-of-the-art graph libraries shows its good performance. Finally, we study the crucial graph algorithm, transitive closure. We propose a disk-based transitive closure solution that operates on the bit-matrix. The solution is within the Python ecosystem while adhering to the principles of database systems. Our experimental study shows the superiority of our solutions over existing popular analysis systems, suggesting potential advancements in bridging high-performance computing and Python."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/10657/17736"],"dc:language.iso":["en"],"dc:subject":["Graph analytics","Big data","Graph metrics","Graph patterns"],"dc:title":["Graph Analytics with Data Science Languages"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Doctor of Philosophy"],"thesis:institution_name":["University of Houston"]},"updated_at":"2026-07-24T02:32:54Z"}