{"id":{"repo_id":"umkc","oai_identifier":"oai:mospace.umsystem.edu:10355/67033"},"canonical_url":"https://search.dev.ndltd.org/etd/umkc/oai:mospace.umsystem.edu:10355/67033","repository":{"repo_id":"umkc","name":"University of Missouri - Kansas City","base_url":"https://mospace.umsystem.edu/oai/request"},"display":{"title":"An Approach for Fast Score Computation in Bayesian Network Structure Learning Over Large-Scale Distributed Data","abstract":"On a fundamental level, Bayes’ theorem enables us to utilize prior knowledge to determine the probability of an event. In consequence, its suitability for probabilistic reasoning has lead to its employment in probabilistic graphical modeling and the inception of Bayesian networks. The ﬁeld is saturated with techniques to learn the structure of a Bayesian network (also known as Bayes network). Nevertheless, most of the techniques struggle when the number of variables (or network nodes) and the input data grow drastically. At that point, parallel distributed processing is the best alternative to alleviate the computational complexity of this problem. To this end, we propose a gossip-based distributed score computation approach called DiSC that is used to compute the suﬃcient statistics of families of variables in order to accelerate the structure learning process of Bayesian networks. We show that DiSC can signiﬁcantly outperform map-reduce style score computations executed by the distributed computation framework Apache Spark on a variety of synthetic and real datasets with a low accuracy trade-oﬀ.","abstract_html":"On a fundamental level, Bayes’ theorem enables us to utilize prior knowledge to determine the probability of an event. In consequence, its suitability for probabilistic reasoning has lead to its employment in probabilistic graphical modeling and the inception of Bayesian networks. The ﬁeld is saturated with techniques to learn the structure of a Bayesian network (also known as Bayes network). Nevertheless, most of the techniques struggle when the number of variables (or network nodes) and the input data grow drastically. At that point, parallel distributed processing is the best alternative to alleviate the computational complexity of this problem. To this end, we propose a gossip-based distributed score computation approach called DiSC that is used to compute the suﬃcient statistics of families of variables in order to accelerate the structure learning process of Bayesian networks. We show that DiSC can signiﬁcantly outperform map-reduce style score computations executed by the distributed computation framework Apache Spark on a variety of synthetic and real datasets with a low accuracy trade-oﬀ.","abstract_has_math":false,"creators":["Katib, Anas Adnan"],"institution":"University of Missouri -- Kansas City","degree_name":"Ph.D.","degree_level":"Doctoral","degree_discipline":"Computer Science, Computer Networking and Communication Systems (UMKC)","degree_department":null,"school":null,"contributors":[],"advisors":["Rao, Praveen R."],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018","date_published":"2018","updated_at":"2026-07-24T05:19:28Z","subjects":[],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10355/67033","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Rao, Praveen R."]},{"key":"dc:creator","label":"Author","values":["Katib, Anas Adnan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2019-01-29T18:35:34Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2019-01-29T18:35:34Z"]},{"key":"dc:date.issued","label":"Date","values":["2018"]},{"key":"dc:publisher","label":"Institution","values":["University of Missouri -- Kansas City"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science, Computer Networking and Communication Systems (UMKC)"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Doctoral"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Missouri--Kansas City"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10355/67033"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Title from PDF of title page viewed February 4, 2019","Dissertation advisor: Praveen Rao","Vita","Includes bibliographical references (pages 53-58)","Thesis (Ph.D.)--School of Computing and Engineering. University of Missouri--Kansas City, 2018"]},{"key":"dc:description.abstract","label":"Abstract","values":["On a fundamental level, Bayes’ theorem enables us to utilize prior knowledge to determine the probability of an event. In consequence, its suitability for probabilistic reasoning has lead to its employment in probabilistic graphical modeling and the inception of Bayesian networks. The ﬁeld is saturated with techniques to learn the structure of a Bayesian network (also known as Bayes network). Nevertheless, most of the techniques struggle when the number of variables (or network nodes) and the input data grow drastically. At that point, parallel distributed processing is the best alternative to alleviate the computational complexity of this problem. To this end, we propose a gossip-based distributed score computation approach called DiSC that is used to compute the suﬃcient statistics of families of variables in order to accelerate the structure learning process of Bayesian networks. We show that DiSC can signiﬁcantly outperform map-reduce style score computations executed by the distributed computation framework Apache Spark on a variety of synthetic and real datasets with a low accuracy trade-oﬀ."]},{"key":"dc:title","label":"Title","values":["An Approach for Fast Score Computation in Bayesian Network Structure Learning Over Large-Scale Distributed Data"]}]}],"canonical_facts":{"dc:contributor.advisor":["Rao, Praveen R."],"dc:creator":["Katib, Anas Adnan"],"dc:date.accessioned":["2019-01-29T18:35:34Z"],"dc:date.available":["2019-01-29T18:35:34Z"],"dc:date.issued":["2018"],"dc:description":["Title from PDF of title page viewed February 4, 2019","Dissertation advisor: Praveen Rao","Vita","Includes bibliographical references (pages 53-58)","Thesis (Ph.D.)--School of Computing and Engineering. University of Missouri--Kansas City, 2018"],"dc:description.abstract":["On a fundamental level, Bayes’ theorem enables us to utilize prior knowledge to determine the probability of an event. In consequence, its suitability for probabilistic reasoning has lead to its employment in probabilistic graphical modeling and the inception of Bayesian networks. The ﬁeld is saturated with techniques to learn the structure of a Bayesian network (also known as Bayes network). Nevertheless, most of the techniques struggle when the number of variables (or network nodes) and the input data grow drastically. At that point, parallel distributed processing is the best alternative to alleviate the computational complexity of this problem. To this end, we propose a gossip-based distributed score computation approach called DiSC that is used to compute the suﬃcient statistics of families of variables in order to accelerate the structure learning process of Bayesian networks. We show that DiSC can signiﬁcantly outperform map-reduce style score computations executed by the distributed computation framework Apache Spark on a variety of synthetic and real datasets with a low accuracy trade-oﬀ."],"dc:identifier.uri":["https://hdl.handle.net/10355/67033"],"dc:language.iso":["en_US"],"dc:publisher":["University of Missouri -- Kansas City"],"dc:title":["An Approach for Fast Score Computation in Bayesian Network Structure Learning Over Large-Scale Distributed Data"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science, Computer Networking and Communication Systems (UMKC)"],"thesis:degree_level":["Doctoral"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Missouri--Kansas City"]},"updated_at":"2026-07-24T05:19:28Z"}