{"id":{"repo_id":"arkansas","oai_identifier":"oai:scholarworks.uark.edu:etd-6154"},"canonical_url":"https://search.dev.ndltd.org/etd/arkansas/oai:scholarworks.uark.edu:etd-6154","repository":{"repo_id":"arkansas","name":"University of Arkansas","base_url":"https://scholarworks.uark.edu/do/oai/"},"display":{"title":"Effective Knowledge Graph Aggregation for Malware-Related Cybersecurity Text","abstract":"<p>With the rate at which malware spreads in the modern age, it is extremely important that cyber security analysts are able to extract relevant information pertaining to new and active threats in a timely and effective manner. Having to manually read through articles and blog posts on the internet is time consuming and usually involves sifting through much repeated information. Knowledge graphs, a structured representation of relationship information, are an effective way to visually condense information presented in large amounts of unstructured text for human readers. Thusly, they are useful for sifting through the abundance of cyber security information that is released through web-based security articles and blogs. This paper presents a pipeline for extracting these relationships using supervised deep learning with the recent state-of-the-art transformer-based neural architectures for sequence processing tasks. To this end, a corpus of text from a range of prominent cybersecurity-focused media outlets was manually annotated. An algorithm is also presented that keeps potentially redundant relationships from being added to an existing knowledge graph, using a cosine-similarity metric on pre-trained word embeddings. </p>","abstract_html":"&lt;p&gt;With the rate at which malware spreads in the modern age, it is extremely important that cyber security analysts are able to extract relevant information pertaining to new and active threats in a timely and effective manner. Having to manually read through articles and blog posts on the internet is time consuming and usually involves sifting through much repeated information. Knowledge graphs, a structured representation of relationship information, are an effective way to visually condense information presented in large amounts of unstructured text for human readers. Thusly, they are useful for sifting through the abundance of cyber security information that is released through web-based security articles and blogs. This paper presents a pipeline for extracting these relationships using supervised deep learning with the recent state-of-the-art transformer-based neural architectures for sequence processing tasks. To this end, a corpus of text from a range of prominent cybersecurity-focused media outlets was manually annotated. An algorithm is also presented that keeps potentially redundant relationships from being added to an existing knowledge graph, using a cosine-similarity metric on pre-trained word embeddings. &lt;/p&gt;","abstract_has_math":false,"creators":["Boudreau, Phillip Ryan"],"institution":null,"degree_name":"Master of Science in Computer Science (MS)","degree_level":"Thesis","degree_discipline":null,"degree_department":null,"school":null,"contributors":["Thompson, Dale R.","Panda, Brajendra N."],"advisors":["Li, Qinghua"],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-08-01T07:00:00Z","date_published":"2022-08-01T07:00:00Z","updated_at":"2026-07-24T00:58:53Z","subjects":["Cyber security","Knowledge graphs","Extraction language models","Information Security","Programming Languages and Compilers"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarworks.uark.edu/etd/4604","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Thompson, Dale R.","Panda, Brajendra N."]},{"key":"dc:contributor.advisor","label":"Advisor","values":["Li, Qinghua"]},{"key":"dc:creator","label":"Author","values":["Boudreau, Phillip Ryan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-02-06T08:00:00Z"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science in Computer Science (MS)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Cyber security","Knowledge graphs","Extraction language models","Information Security","Programming Languages and Compilers"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarworks.uark.edu/etd/4604"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>With the rate at which malware spreads in the modern age, it is extremely important that cyber security analysts are able to extract relevant information pertaining to new and active threats in a timely and effective manner. Having to manually read through articles and blog posts on the internet is time consuming and usually involves sifting through much repeated information. Knowledge graphs, a structured representation of relationship information, are an effective way to visually condense information presented in large amounts of unstructured text for human readers. Thusly, they are useful for sifting through the abundance of cyber security information that is released through web-based security articles and blogs. This paper presents a pipeline for extracting these relationships using supervised deep learning with the recent state-of-the-art transformer-based neural architectures for sequence processing tasks. To this end, a corpus of text from a range of prominent cybersecurity-focused media outlets was manually annotated. An algorithm is also presented that keeps potentially redundant relationships from being added to an existing knowledge graph, using a cosine-similarity metric on pre-trained word embeddings. </p>"]},{"key":"dc:title","label":"Title","values":["Effective Knowledge Graph Aggregation for Malware-Related Cybersecurity Text"]}]}],"canonical_facts":{"dc:contributor":["Thompson, Dale R.","Panda, Brajendra N."],"dc:contributor.advisor":["Li, Qinghua"],"dc:creator":["Boudreau, Phillip Ryan"],"dc:date":["2022"],"dc:date.available":["2024-02-06T08:00:00Z"],"dc:description.abstract":["<p>With the rate at which malware spreads in the modern age, it is extremely important that cyber security analysts are able to extract relevant information pertaining to new and active threats in a timely and effective manner. Having to manually read through articles and blog posts on the internet is time consuming and usually involves sifting through much repeated information. Knowledge graphs, a structured representation of relationship information, are an effective way to visually condense information presented in large amounts of unstructured text for human readers. Thusly, they are useful for sifting through the abundance of cyber security information that is released through web-based security articles and blogs. This paper presents a pipeline for extracting these relationships using supervised deep learning with the recent state-of-the-art transformer-based neural architectures for sequence processing tasks. To this end, a corpus of text from a range of prominent cybersecurity-focused media outlets was manually annotated. An algorithm is also presented that keeps potentially redundant relationships from being added to an existing knowledge graph, using a cosine-similarity metric on pre-trained word embeddings. </p>"],"dc:identifier":["https://scholarworks.uark.edu/etd/4604"],"dc:subject":["Cyber security","Knowledge graphs","Extraction language models","Information Security","Programming Languages and Compilers"],"dc:title":["Effective Knowledge Graph Aggregation for Malware-Related Cybersecurity Text"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["Master of Science in Computer Science (MS)"]},"updated_at":"2026-07-24T00:58:53Z"}