{"id":{"repo_id":"embry-riddle","oai_identifier":"oai:commons.erau.edu:edt-1813"},"canonical_url":"https://search.dev.ndltd.org/etd/embry-riddle/oai:commons.erau.edu:edt-1813","repository":{"repo_id":"embry-riddle","name":"Embry Riddle Aeronautical University","base_url":"https://commons.erau.edu/do/oai/"},"display":{"title":"Spoken Language Processing and Modeling for Aviation Communications","abstract":"<p>With recent advances in machine learning and deep learning technologies and the creation of larger aviation-specific corpora, applying natural language processing technologies, especially those based on transformer neural networks, to aviation communications is becoming increasingly feasible. Previous work has focused on machine learning applications to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained checkpoints and those trained from their base weight initializations (trained from scratch). The results suggest that transformer language models trained from scratch outperform models fine-tuned from pretrained checkpoints. The work concludes by recommending future work to improve pretraining performance and suggestions for downstream, in-domain tasks such as semantic extraction, named entity recognition (callsign identification), speaker role identification, and speech recognition.</p>","abstract_html":"&lt;p&gt;With recent advances in machine learning and deep learning technologies and the creation of larger aviation-specific corpora, applying natural language processing technologies, especially those based on transformer neural networks, to aviation communications is becoming increasingly feasible. Previous work has focused on machine learning applications to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained checkpoints and those trained from their base weight initializations (trained from scratch). The results suggest that transformer language models trained from scratch outperform models fine-tuned from pretrained checkpoints. The work concludes by recommending future work to improve pretraining performance and suggestions for downstream, in-domain tasks such as semantic extraction, named entity recognition (callsign identification), speaker role identification, and speech recognition.&lt;/p&gt;","abstract_has_math":false,"creators":["Van De Brook, Aaron"],"institution":null,"degree_name":"Master of Science in Electrical & Computer Engineering","degree_level":"Thesis - Open Access","degree_discipline":"Electrical Engineering and Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-10-01T07:00:00Z","date_published":"2023-10-01T07:00:00Z","updated_at":"2026-07-27T19:26:16Z","subjects":["language modeling","automatic speech recognition","corpus processing","aviation communications","deep learning","machine learning","Artificial Intelligence and Robotics","Data Science","Signal Processing","Systems and Communications"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://commons.erau.edu/edt/788","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Van De Brook, Aaron"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Electrical Engineering and Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis - Open Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science in Electrical & Computer Engineering"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["language modeling","automatic speech recognition","corpus processing","aviation communications","deep learning","machine learning","Artificial Intelligence and Robotics","Data Science","Signal Processing","Systems and Communications"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://commons.erau.edu/edt/788"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>With recent advances in machine learning and deep learning technologies and the creation of larger aviation-specific corpora, applying natural language processing technologies, especially those based on transformer neural networks, to aviation communications is becoming increasingly feasible. Previous work has focused on machine learning applications to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained checkpoints and those trained from their base weight initializations (trained from scratch). The results suggest that transformer language models trained from scratch outperform models fine-tuned from pretrained checkpoints. The work concludes by recommending future work to improve pretraining performance and suggestions for downstream, in-domain tasks such as semantic extraction, named entity recognition (callsign identification), speaker role identification, and speech recognition.</p>"]},{"key":"dc:title","label":"Title","values":["Spoken Language Processing and Modeling for Aviation Communications"]}]}],"canonical_facts":{"dc:creator":["Van De Brook, Aaron"],"dc:description.abstract":["<p>With recent advances in machine learning and deep learning technologies and the creation of larger aviation-specific corpora, applying natural language processing technologies, especially those based on transformer neural networks, to aviation communications is becoming increasingly feasible. Previous work has focused on machine learning applications to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained checkpoints and those trained from their base weight initializations (trained from scratch). The results suggest that transformer language models trained from scratch outperform models fine-tuned from pretrained checkpoints. The work concludes by recommending future work to improve pretraining performance and suggestions for downstream, in-domain tasks such as semantic extraction, named entity recognition (callsign identification), speaker role identification, and speech recognition.</p>"],"dc:identifier":["https://commons.erau.edu/edt/788"],"dc:subject":["language modeling","automatic speech recognition","corpus processing","aviation communications","deep learning","machine learning","Artificial Intelligence and Robotics","Data Science","Signal Processing","Systems and Communications"],"dc:title":["Spoken Language Processing and Modeling for Aviation Communications"],"thesis:degree_discipline":["Electrical Engineering and Computer Science"],"thesis:degree_level":["Thesis - Open Access"],"thesis:degree_name":["Master of Science in Electrical & Computer Engineering"]},"updated_at":"2026-07-27T19:26:16Z"}