{"id":{"repo_id":"umkc","oai_identifier":"oai:mospace.umsystem.edu:10355/84193"},"canonical_url":"https://search.dev.ndltd.org/etd/umkc/oai:mospace.umsystem.edu:10355/84193","repository":{"repo_id":"umkc","name":"University of Missouri - Kansas City","base_url":"https://mospace.umsystem.edu/oai/request"},"display":{"title":"Shared Context through Multi-Level Attention Transformers for Text Classification","abstract":"Natural language processing (NLP) has seen recent explosive growth by creating artificial intelligence with human-level intelligence. Understanding the context using an attention mechanism could be further improved by fine-tuning their composition for classification, question answering, and topic modeling. Real-world datasets are much more complex and tend to require multi-fold models. Such models tend to be larger, deeper, more complicated; for example, BERT has 340 million parameters, Turing NLG is 17billion parameters, and GPT-3 is about 175 billion parameters. Understanding their implications requires the immense computational ability to process the text corpus during both training and inferences. This thesis proposes a novel deep learning architecture for scalable multi-fold text classification that is an extension of BERT by sharing context across abstraction levels of domains. Four types of deep learning models (BERT flat, BERT hierarchical, BERT hierarchical tuned, BERT Feature extracted) are proposed for the multi-label attention transformers on the architecture. The proposed models provide a means to overcome competing limitations, training concurrently, and providing predictions for an extra level of classes simultaneously. Our work overcomes the limitations of knowledge distillation or transfer-learning, i.e., it is not scalable or sustainable, and it’s also costly. We have performed experiments to validate the reliability model using both benchmark and real-world data (KCMO 311 data). Quantitative results confirm that the proposed models can enhance model performance in terms of computational requirements and provide competitive accuracy.","abstract_html":"Natural language processing (NLP) has seen recent explosive growth by creating artificial intelligence with human-level intelligence. Understanding the context using an attention mechanism could be further improved by fine-tuning their composition for classification, question answering, and topic modeling. Real-world datasets are much more complex and tend to require multi-fold models. Such models tend to be larger, deeper, more complicated; for example, BERT has 340 million parameters, Turing NLG is 17billion parameters, and GPT-3 is about 175 billion parameters. Understanding their implications requires the immense computational ability to process the text corpus during both training and inferences. This thesis proposes a novel deep learning architecture for scalable multi-fold text classification that is an extension of BERT by sharing context across abstraction levels of domains. Four types of deep learning models (BERT flat, BERT hierarchical, BERT hierarchical tuned, BERT Feature extracted) are proposed for the multi-label attention transformers on the architecture. The proposed models provide a means to overcome competing limitations, training concurrently, and providing predictions for an extra level of classes simultaneously. Our work overcomes the limitations of knowledge distillation or transfer-learning, i.e., it is not scalable or sustainable, and it’s also costly. We have performed experiments to validate the reliability model using both benchmark and real-world data (KCMO 311 data). Quantitative results confirm that the proposed models can enhance model performance in terms of computational requirements and provide competitive accuracy.","abstract_has_math":false,"creators":["Thota, Charan Tej"],"institution":"University of Missouri--Kansas City","degree_name":"M.S. (Master of Science)","degree_level":"Masters","degree_discipline":"Computer Science (UMKC)","degree_department":null,"school":null,"contributors":[],"advisors":["Lee, Yugyung, 1960-"],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021","date_published":"2021","updated_at":"2026-07-24T05:19:02Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10355/84193","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Lee, Yugyung, 1960-"]},{"key":"dc:creator","label":"Author","values":["Thota, Charan Tej"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2021-06-02T16:13:30Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2021-06-02T16:13:30Z"]},{"key":"dc:date.issued","label":"Date","values":["2021"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science (UMKC)"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S. (Master of Science)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Missouri--Kansas City"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10355/84193"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Title from PDF of title page viewed June 14, 2021","Thesis advisor: Yugyung Lee","Vita","Includes bibliographical references (pages 79-80)","Thesis (M.S.)--School of Computing and Engineering. University of Missouri--Kansas City, 2021"]},{"key":"dc:description.abstract","label":"Abstract","values":["Natural language processing (NLP) has seen recent explosive growth by creating artificial intelligence with human-level intelligence. Understanding the context using an attention mechanism could be further improved by fine-tuning their composition for classification, question answering, and topic modeling. Real-world datasets are much more complex and tend to require multi-fold models. Such models tend to be larger, deeper, more complicated; for example, BERT has 340 million parameters, Turing NLG is 17billion parameters, and GPT-3 is about 175 billion parameters. Understanding their implications requires the immense computational ability to process the text corpus during both training and inferences. This thesis proposes a novel deep learning architecture for scalable multi-fold text classification that is an extension of BERT by sharing context across abstraction levels of domains. Four types of deep learning models (BERT flat, BERT hierarchical, BERT hierarchical tuned, BERT Feature extracted) are proposed for the multi-label attention transformers on the architecture. The proposed models provide a means to overcome competing limitations, training concurrently, and providing predictions for an extra level of classes simultaneously. Our work overcomes the limitations of knowledge distillation or transfer-learning, i.e., it is not scalable or sustainable, and it’s also costly. We have performed experiments to validate the reliability model using both benchmark and real-world data (KCMO 311 data). Quantitative results confirm that the proposed models can enhance model performance in terms of computational requirements and provide competitive accuracy."]},{"key":"dc:title","label":"Title","values":["Shared Context through Multi-Level Attention Transformers for Text Classification"]}]}],"canonical_facts":{"dc:contributor.advisor":["Lee, Yugyung, 1960-"],"dc:creator":["Thota, Charan Tej"],"dc:date.accessioned":["2021-06-02T16:13:30Z"],"dc:date.available":["2021-06-02T16:13:30Z"],"dc:date.issued":["2021"],"dc:description":["Title from PDF of title page viewed June 14, 2021","Thesis advisor: Yugyung Lee","Vita","Includes bibliographical references (pages 79-80)","Thesis (M.S.)--School of Computing and Engineering. University of Missouri--Kansas City, 2021"],"dc:description.abstract":["Natural language processing (NLP) has seen recent explosive growth by creating artificial intelligence with human-level intelligence. Understanding the context using an attention mechanism could be further improved by fine-tuning their composition for classification, question answering, and topic modeling. Real-world datasets are much more complex and tend to require multi-fold models. Such models tend to be larger, deeper, more complicated; for example, BERT has 340 million parameters, Turing NLG is 17billion parameters, and GPT-3 is about 175 billion parameters. Understanding their implications requires the immense computational ability to process the text corpus during both training and inferences. This thesis proposes a novel deep learning architecture for scalable multi-fold text classification that is an extension of BERT by sharing context across abstraction levels of domains. Four types of deep learning models (BERT flat, BERT hierarchical, BERT hierarchical tuned, BERT Feature extracted) are proposed for the multi-label attention transformers on the architecture. The proposed models provide a means to overcome competing limitations, training concurrently, and providing predictions for an extra level of classes simultaneously. Our work overcomes the limitations of knowledge distillation or transfer-learning, i.e., it is not scalable or sustainable, and it’s also costly. We have performed experiments to validate the reliability model using both benchmark and real-world data (KCMO 311 data). Quantitative results confirm that the proposed models can enhance model performance in terms of computational requirements and provide competitive accuracy."],"dc:identifier.uri":["https://hdl.handle.net/10355/84193"],"dc:title":["Shared Context through Multi-Level Attention Transformers for Text Classification"],"thesis:degree_discipline":["Computer Science (UMKC)"],"thesis:degree_level":["Masters"],"thesis:degree_name":["M.S. (Master of Science)"],"thesis:institution_name":["University of Missouri--Kansas City"]},"updated_at":"2026-07-24T05:19:02Z"}