{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/106431"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/106431","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Machine learning-based analytics of structured and unstructured data for enhanced bridge deterioration prediction","abstract":"The increasing availability of heterogeneous bridge data from multiple sources opens unprecedented opportunities for data analytics to better predict bridge deterioration for supporting enhanced bridge maintenance decision making. Such data include structured National Bridge Inventory (NBI) and National Bridge Elements (NBE) data, structured traffic and weather data, and unstructured textual bridge inspection reports. However, despite the availability of the data, existing data-driven prediction methods mostly learn from abstract inventory data (e.g., the NBI data which describe bridge conditions by condition ratings) from a single source – missing the opportunity of leveraging the wealth of unstructured textual inspection reports and the diverseness of the multi-source data for enhanced deterioration prediction. To capitalize on this opportunity, a novel bridge data analytics framework is proposed. The proposed framework is composed of six primary components: (1) a bridge deterioration knowledge ontology for facilitating semantic information and relation extraction from textual bridge inspection reports based on content and domain-specific meaning; (2) a semi-supervised machine learning-based semantic information extraction method for extracting information entities that describe bridge conditions and maintenance actions from the reports; (3) a supervised machine learning-based semantic relation extraction method for extracting dependency relations from the reports to link the extracted, yet isolated, information entities into concepts and to represent the semantically-low concepts in a semantically-rich structured way; (4) an unsupervised machine learning-based data linking method for linking the data records that are extracted from the reports and refer to the same entity; (5) a hybrid data fusion method for fusing the linked data records into a unified representation and for, subsequently, integrating the fused data with the other types of structured data (i.e., NBI and NBE data, as well as traffic and weather data); and (6) a data-driven, deep learning-based bridge deterioration prediction method for learning from the integrated bridge data to predict the condition ratings of the primary bridge components and to predict the quantities of specific bridge element-level deficiencies. The performance of the proposed framework was evaluated in predicting the deterioration of the state-owned bridges in Washington. It achieved a macro-precision and macro-recall of 89.9% and 85.8% when predicting the future condition ratings of the primary bridge components (i.e., decks, superstructures, and substructures), and achieved a root mean square error, coefficient of variation, and coefficient of determination of 1.3, 27.6%, and 0.89, respectively, when predicting the future quantities of specific bridge element-level deficiencies. The experimental results demonstrated the promise of the proposed framework.","abstract_html":"The increasing availability of heterogeneous bridge data from multiple sources opens unprecedented opportunities for data analytics to better predict bridge deterioration for supporting enhanced bridge maintenance decision making. Such data include structured National Bridge Inventory (NBI) and National Bridge Elements (NBE) data, structured traffic and weather data, and unstructured textual bridge inspection reports. However, despite the availability of the data, existing data-driven prediction methods mostly learn from abstract inventory data (e.g., the NBI data which describe bridge conditions by condition ratings) from a single source – missing the opportunity of leveraging the wealth of unstructured textual inspection reports and the diverseness of the multi-source data for enhanced deterioration prediction. To capitalize on this opportunity, a novel bridge data analytics framework is proposed. The proposed framework is composed of six primary components: (1) a bridge deterioration knowledge ontology for facilitating semantic information and relation extraction from textual bridge inspection reports based on content and domain-specific meaning; (2) a semi-supervised machine learning-based semantic information extraction method for extracting information entities that describe bridge conditions and maintenance actions from the reports; (3) a supervised machine learning-based semantic relation extraction method for extracting dependency relations from the reports to link the extracted, yet isolated, information entities into concepts and to represent the semantically-low concepts in a semantically-rich structured way; (4) an unsupervised machine learning-based data linking method for linking the data records that are extracted from the reports and refer to the same entity; (5) a hybrid data fusion method for fusing the linked data records into a unified representation and for, subsequently, integrating the fused data with the other types of structured data (i.e., NBI and NBE data, as well as traffic and weather data); and (6) a data-driven, deep learning-based bridge deterioration prediction method for learning from the integrated bridge data to predict the condition ratings of the primary bridge components and to predict the quantities of specific bridge element-level deficiencies. The performance of the proposed framework was evaluated in predicting the deterioration of the state-owned bridges in Washington. It achieved a macro-precision and macro-recall of 89.9% and 85.8% when predicting the future condition ratings of the primary bridge components (i.e., decks, superstructures, and substructures), and achieved a root mean square error, coefficient of variation, and coefficient of determination of 1.3, 27.6%, and 0.89, respectively, when predicting the future quantities of specific bridge element-level deficiencies. The experimental results demonstrated the promise of the proposed framework.","abstract_has_math":false,"creators":["Liu, Kaijian"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Civil Engineering","degree_department":null,"school":null,"contributors":["El-Gohary, Nora","El-Rayes, Khaled","Zhai, ChengXiang","Liu, Liang Y","Golparvar-Fard, Mani"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-03-02T22:38:38Z","date_published":"2020-03-02T22:38:38Z","updated_at":"2026-07-22T22:24:47Z","subjects":["Bridge deterioration prediction","Data analytics","Machine learning","Natural language processing","Ontology","Information extraction","Dependency parsing","Data linking","Data fusion"],"languages":["en"],"rights":["Copyright 2019 Kaijian Liu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/106431","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["El-Gohary, Nora","El-Rayes, Khaled","Zhai, ChengXiang","Liu, Liang Y","Golparvar-Fard, Mani"]},{"key":"dc:creator","label":"Author","values":["Liu, Kaijian"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2020-03-02T22:38:38Z","2022-03-03T10:15:22Z","2019-10-03","2019-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Civil Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bridge deterioration prediction","Data analytics","Machine learning","Natural language processing","Ontology","Information extraction","Dependency parsing","Data linking","Data fusion"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2019 Kaijian Liu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/106431"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The increasing availability of heterogeneous bridge data from multiple sources opens unprecedented opportunities for data analytics to better predict bridge deterioration for supporting enhanced bridge maintenance decision making. Such data include structured National Bridge Inventory (NBI) and National Bridge Elements (NBE) data, structured traffic and weather data, and unstructured textual bridge inspection reports. However, despite the availability of the data, existing data-driven prediction methods mostly learn from abstract inventory data (e.g., the NBI data which describe bridge conditions by condition ratings) from a single source – missing the opportunity of leveraging the wealth of unstructured textual inspection reports and the diverseness of the multi-source data for enhanced deterioration prediction. To capitalize on this opportunity, a novel bridge data analytics framework is proposed. The proposed framework is composed of six primary components: (1) a bridge deterioration knowledge ontology for facilitating semantic information and relation extraction from textual bridge inspection reports based on content and domain-specific meaning; (2) a semi-supervised machine learning-based semantic information extraction method for extracting information entities that describe bridge conditions and maintenance actions from the reports; (3) a supervised machine learning-based semantic relation extraction method for extracting dependency relations from the reports to link the extracted, yet isolated, information entities into concepts and to represent the semantically-low concepts in a semantically-rich structured way; (4) an unsupervised machine learning-based data linking method for linking the data records that are extracted from the reports and refer to the same entity; (5) a hybrid data fusion method for fusing the linked data records into a unified representation and for, subsequently, integrating the fused data with the other types of structured data (i.e., NBI and NBE data, as well as traffic and weather data); and (6) a data-driven, deep learning-based bridge deterioration prediction method for learning from the integrated bridge data to predict the condition ratings of the primary bridge components and to predict the quantities of specific bridge element-level deficiencies. The performance of the proposed framework was evaluated in predicting the deterioration of the state-owned bridges in Washington. It achieved a macro-precision and macro-recall of 89.9% and 85.8% when predicting the future condition ratings of the primary bridge components (i.e., decks, superstructures, and substructures), and achieved a root mean square error, coefficient of variation, and coefficient of determination of 1.3, 27.6%, and 0.89, respectively, when predicting the future quantities of specific bridge element-level deficiencies. The experimental results demonstrated the promise of the proposed framework.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-12-01","The student, Kaijian Liu, accepted the attached license on 2019-10-02 at 09:48.","The student, Kaijian Liu, submitted this Dissertation for approval on 2019-10-02 at 10:13.","This Dissertation was approved for publication on 2019-10-03 at 10:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14482 on 2020-02-28 at 17:35:38","Made available in DSpace on 2020-03-02T22:38:38Z (GMT). No. of bitstreams: 2 LIU-DISSERTATION-2019.pdf: 4762751 bytes, checksum: 59829c227b2ac108e2558c20b7252150 (MD5) LICENSE.txt: 4208 bytes, checksum: 21fedf92e10ac12a33a97aae13808cc6 (MD5) Previous issue date: 2019-10-03","Embargo set by: Seth Robbins for item 113975 Lift date: 2022-03-02T22:39:04Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113975 on 2022-03-03T10:15:22Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Machine learning-based analytics of structured and unstructured data for enhanced bridge deterioration prediction"]}]}],"canonical_facts":{"dc:contributor":["El-Gohary, Nora","El-Rayes, Khaled","Zhai, ChengXiang","Liu, Liang Y","Golparvar-Fard, Mani"],"dc:creator":["Liu, Kaijian"],"dc:date":["2020-03-02T22:38:38Z","2022-03-03T10:15:22Z","2019-10-03","2019-12"],"dc:description":["The increasing availability of heterogeneous bridge data from multiple sources opens unprecedented opportunities for data analytics to better predict bridge deterioration for supporting enhanced bridge maintenance decision making. Such data include structured National Bridge Inventory (NBI) and National Bridge Elements (NBE) data, structured traffic and weather data, and unstructured textual bridge inspection reports. However, despite the availability of the data, existing data-driven prediction methods mostly learn from abstract inventory data (e.g., the NBI data which describe bridge conditions by condition ratings) from a single source – missing the opportunity of leveraging the wealth of unstructured textual inspection reports and the diverseness of the multi-source data for enhanced deterioration prediction. To capitalize on this opportunity, a novel bridge data analytics framework is proposed. The proposed framework is composed of six primary components: (1) a bridge deterioration knowledge ontology for facilitating semantic information and relation extraction from textual bridge inspection reports based on content and domain-specific meaning; (2) a semi-supervised machine learning-based semantic information extraction method for extracting information entities that describe bridge conditions and maintenance actions from the reports; (3) a supervised machine learning-based semantic relation extraction method for extracting dependency relations from the reports to link the extracted, yet isolated, information entities into concepts and to represent the semantically-low concepts in a semantically-rich structured way; (4) an unsupervised machine learning-based data linking method for linking the data records that are extracted from the reports and refer to the same entity; (5) a hybrid data fusion method for fusing the linked data records into a unified representation and for, subsequently, integrating the fused data with the other types of structured data (i.e., NBI and NBE data, as well as traffic and weather data); and (6) a data-driven, deep learning-based bridge deterioration prediction method for learning from the integrated bridge data to predict the condition ratings of the primary bridge components and to predict the quantities of specific bridge element-level deficiencies. The performance of the proposed framework was evaluated in predicting the deterioration of the state-owned bridges in Washington. It achieved a macro-precision and macro-recall of 89.9% and 85.8% when predicting the future condition ratings of the primary bridge components (i.e., decks, superstructures, and substructures), and achieved a root mean square error, coefficient of variation, and coefficient of determination of 1.3, 27.6%, and 0.89, respectively, when predicting the future quantities of specific bridge element-level deficiencies. The experimental results demonstrated the promise of the proposed framework.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2021-12-01","The student, Kaijian Liu, accepted the attached license on 2019-10-02 at 09:48.","The student, Kaijian Liu, submitted this Dissertation for approval on 2019-10-02 at 10:13.","This Dissertation was approved for publication on 2019-10-03 at 10:11.","DSpace SAF Submission Ingestion Package generated from Vireo submission #14482 on 2020-02-28 at 17:35:38","Made available in DSpace on 2020-03-02T22:38:38Z (GMT). No. of bitstreams: 2 LIU-DISSERTATION-2019.pdf: 4762751 bytes, checksum: 59829c227b2ac108e2558c20b7252150 (MD5) LICENSE.txt: 4208 bytes, checksum: 21fedf92e10ac12a33a97aae13808cc6 (MD5) Previous issue date: 2019-10-03","Embargo set by: Seth Robbins for item 113975 Lift date: 2022-03-02T22:39:04Z Reason: Author requested closed access (OA after 2yrs) in Vireo ETD system","Limited Restriction Lifted for Item 113975 on 2022-03-03T10:15:22Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/106431"],"dc:language":["en"],"dc:rights":["Copyright 2019 Kaijian Liu"],"dc:subject":["Bridge deterioration prediction","Data analytics","Machine learning","Natural language processing","Ontology","Information extraction","Dependency parsing","Data linking","Data fusion"],"dc:title":["Machine learning-based analytics of structured and unstructured data for enhanced bridge deterioration prediction"],"dc:type":["text"],"thesis:degree_discipline":["Civil Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:47Z"}