{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/385967"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/385967","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Text and data mining of semiconductor material information from scientific literature","abstract":"The scientific literature holds vast amount of unstructured information on material properties. Given the right toolkits, material property data can be extracted from text with little user interaction, and analysed for patterns. This data-driven approach is the first step towards accelerated material discovery. This thesis focuses on the development and application of such text and data mining software toolkits for semiconductor band gap information, using natural language processing techniques, machine-learning algorithms, and language models. Chapter 1 gives a description of data-driven material discovery, an overview of chemical data extraction from scientific documents, and the motivation of generating semiconductor band gap databases. Chapter 2 presents an automatically generated database of 100,236 semiconductor band gap records, with associated temperature values, via text and data mining on research papers with ChemDataExtractor. Chapter 3 introduces Snowball 2.0, a generic sentence-level parser for chemical data extraction with ChemDataExtractor. It features improved performance, better generalizability, enhanced functionalities, and simpler interaction with users. A Snowball model that was trained and evaluated with semiconductor band gap information is also provided. Chapter 4 presents SemiconductorBERT, a set of transformer-based language models that were trained and optimised to extract band gap information of chemicals from text, with better performance than other openly-available models. Chapter 5 concludes this thesis, and outlines possible directions for future research.","abstract_html":"The scientific literature holds vast amount of unstructured information on material properties. Given the right toolkits, material property data can be extracted from text with little user interaction, and analysed for patterns. This data-driven approach is the first step towards accelerated material discovery. This thesis focuses on the development and application of such text and data mining software toolkits for semiconductor band gap information, using natural language processing techniques, machine-learning algorithms, and language models. Chapter 1 gives a description of data-driven material discovery, an overview of chemical data extraction from scientific documents, and the motivation of generating semiconductor band gap databases. Chapter 2 presents an automatically generated database of 100,236 semiconductor band gap records, with associated temperature values, via text and data mining on research papers with ChemDataExtractor. Chapter 3 introduces Snowball 2.0, a generic sentence-level parser for chemical data extraction with ChemDataExtractor. It features improved performance, better generalizability, enhanced functionalities, and simpler interaction with users. A Snowball model that was trained and evaluated with semiconductor band gap information is also provided. Chapter 4 presents SemiconductorBERT, a set of transformer-based language models that were trained and optimised to extract band gap information of chemicals from text, with better performance than other openly-available models. Chapter 5 concludes this thesis, and outlines possible directions for future research.","abstract_has_math":false,"creators":["Dong, Qingyang"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Jacqueline, Cole"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-16","date_published":"2025-03-16","updated_at":"2026-07-24T01:33:08Z","subjects":["Machine Learning","Natural Language Processing","Python","Semiconductor"],"languages":[],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/0a1e817a-ef80-4068-a8d6-34f72fcb503a/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.119390","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Jacqueline, Cole"]},{"key":"dc:creator","label":"Author","values":["Dong, Qingyang"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-03-16"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/385967"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Machine Learning","Natural Language Processing","Python","Semiconductor"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/0a1e817a-ef80-4068-a8d6-34f72fcb503a/download","http://purl.org/NET/rdflicense/allrightsreserved"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2026-06-24"]},{"key":"dc:rights.embargotype","label":"Dc Rights Embargotype","values":["embargo"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.119390"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/5466c9d4-8ac1-4b70-acee-4e9b7dad8bdd/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The scientific literature holds vast amount of unstructured information on material properties. Given the right toolkits, material property data can be extracted from text with little user interaction, and analysed for patterns. This data-driven approach is the first step towards accelerated material discovery. This thesis focuses on the development and application of such text and data mining software toolkits for semiconductor band gap information, using natural language processing techniques, machine-learning algorithms, and language models. Chapter 1 gives a description of data-driven material discovery, an overview of chemical data extraction from scientific documents, and the motivation of generating semiconductor band gap databases. Chapter 2 presents an automatically generated database of 100,236 semiconductor band gap records, with associated temperature values, via text and data mining on research papers with ChemDataExtractor. Chapter 3 introduces Snowball 2.0, a generic sentence-level parser for chemical data extraction with ChemDataExtractor. It features improved performance, better generalizability, enhanced functionalities, and simpler interaction with users. A Snowball model that was trained and evaluated with semiconductor band gap information is also provided. Chapter 4 presents SemiconductorBERT, a set of transformer-based language models that were trained and optimised to extract band gap information of chemicals from text, with better performance than other openly-available models. Chapter 5 concludes this thesis, and outlines possible directions for future research."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["b4ea112b6c6847ec06fa46642935d3d3","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Text and data mining of semiconductor material information from scientific literature"]}]}],"canonical_facts":{"dc:contributor.advisor":["Jacqueline, Cole"],"dc:creator":["Dong, Qingyang"],"dc:date.issued":["2025-03-16"],"dc:description.abstract":["The scientific literature holds vast amount of unstructured information on material properties. Given the right toolkits, material property data can be extracted from text with little user interaction, and analysed for patterns. This data-driven approach is the first step towards accelerated material discovery. This thesis focuses on the development and application of such text and data mining software toolkits for semiconductor band gap information, using natural language processing techniques, machine-learning algorithms, and language models. Chapter 1 gives a description of data-driven material discovery, an overview of chemical data extraction from scientific documents, and the motivation of generating semiconductor band gap databases. Chapter 2 presents an automatically generated database of 100,236 semiconductor band gap records, with associated temperature values, via text and data mining on research papers with ChemDataExtractor. Chapter 3 introduces Snowball 2.0, a generic sentence-level parser for chemical data extraction with ChemDataExtractor. It features improved performance, better generalizability, enhanced functionalities, and simpler interaction with users. A Snowball model that was trained and evaluated with semiconductor band gap information is also provided. Chapter 4 presents SemiconductorBERT, a set of transformer-based language models that were trained and optimised to extract band gap information of chemicals from text, with better performance than other openly-available models. Chapter 5 concludes this thesis, and outlines possible directions for future research."],"dc:format.checksum.md5":["b4ea112b6c6847ec06fa46642935d3d3","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.119390"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/5466c9d4-8ac1-4b70-acee-4e9b7dad8bdd/download"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/385967"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/0a1e817a-ef80-4068-a8d6-34f72fcb503a/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:rights.embargodate":["2026-06-24"],"dc:rights.embargotype":["embargo"],"dc:subject":["Machine Learning","Natural Language Processing","Python","Semiconductor"],"dc:title":["Text and data mining of semiconductor material information from scientific literature"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T01:33:08Z"}