{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/318837"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/318837","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Automatic creation of novel databases with data-extraction tools for the advancement of third-generation solar cells.","abstract":"This thesis addresses the use of literature-mining software tools to automatically create databases, to accelerate progress in third-generation photovoltaic devices. New discoveries in this area are typically published in research articles, where the data are contained in figures, tables and text. The field could benefit from a single resource where these key data are stored, such that progress might be driven by new machine-learning techniques, applied to large quantities of structured information. The thesis also describes new tools designed to extract other pertinent data. Chapter 1 discusses the potential of third-generation solar devices as alternative fuel sources. Chapter 2 introduces relevant state-of-the-art data-mining techniques and methods used throughout the thesis. Chapter 3 details a bespoke workflow for the extraction of UV/vis absorption spectral peak data from the literature; specifically the wavelength of maximum absorption, λmax, and the molar extinction coefficient, ε. Application of the workflow results in a database of 18,309 chemical records of experimental data. Chapter 4 describes two databases of photovoltaic properties that were auto-generated from the literature, using a data-extraction pipeline that was tailored for this application. The databases contain over 50,000 records of photovoltaic attributes about dye-sensitized solar cells and perovskite solar cells, and all the key properties were extracted with a collective precision of 84%. Chapter 5 introduces ChemSchematicResolver, a high-throughput data-extraction toolkit that can automatically detect chemical-schematic diagrams inside a document, resolve any R-group substituents and convert the resulting diagram to a machine-readable format. The tool was tested on a new evaluation set and all assessed areas achieved precisions of 83-100%. Chapter 6 discusses contributions to two new independent image-extraction software tools, ImageDataExtractor and FigureDataExtractor. ImageDataExtractor is a toolkit designed to extract quantitative data of particle size, shape and distribution from microscopy images. FigureDataExtractor was built to reproduce quantitative data from line graphs. Chapter 7 concludes the work and investigates potential avenues for future research.","abstract_html":"This thesis addresses the use of literature-mining software tools to automatically create databases, to accelerate progress in third-generation photovoltaic devices. New discoveries in this area are typically published in research articles, where the data are contained in figures, tables and text. The field could benefit from a single resource where these key data are stored, such that progress might be driven by new machine-learning techniques, applied to large quantities of structured information. The thesis also describes new tools designed to extract other pertinent data. Chapter 1 discusses the potential of third-generation solar devices as alternative fuel sources. Chapter 2 introduces relevant state-of-the-art data-mining techniques and methods used throughout the thesis. Chapter 3 details a bespoke workflow for the extraction of UV/vis absorption spectral peak data from the literature; specifically the wavelength of maximum absorption, λmax, and the molar extinction coefficient, ε. Application of the workflow results in a database of 18,309 chemical records of experimental data. Chapter 4 describes two databases of photovoltaic properties that were auto-generated from the literature, using a data-extraction pipeline that was tailored for this application. The databases contain over 50,000 records of photovoltaic attributes about dye-sensitized solar cells and perovskite solar cells, and all the key properties were extracted with a collective precision of 84%. Chapter 5 introduces ChemSchematicResolver, a high-throughput data-extraction toolkit that can automatically detect chemical-schematic diagrams inside a document, resolve any R-group substituents and convert the resulting diagram to a machine-readable format. The tool was tested on a new evaluation set and all assessed areas achieved precisions of 83-100%. Chapter 6 discusses contributions to two new independent image-extraction software tools, ImageDataExtractor and FigureDataExtractor. ImageDataExtractor is a toolkit designed to extract quantitative data of particle size, shape and distribution from microscopy images. FigureDataExtractor was built to reproduce quantitative data from line graphs. Chapter 7 concludes the work and investigates potential avenues for future research.","abstract_has_math":false,"creators":["Beard, Edward"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Cole, Jacqueline"],"committee_chairs":[],"committee_members":[],"year":2020,"date_issued":"2020-09-30","date_published":"2020-09-30","updated_at":"2026-07-22T22:24:00Z","subjects":["data mining","dye-sensitized solar cells","perovskite solar cells","text mining","figure mining","UV/vis absorption"],"languages":["eng"],"rights":[],"rights_urls":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/69eb0511-090a-4098-8dd1-de9f4e146b39/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"identifier_entries":[]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.65955","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Cole, Jacqueline"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Tessella Science and Technology Facilities Council (STFC)"]},{"key":"dc:creator","label":"Author","values":["Beard, Edward"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2020-09-30"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/318837"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["data mining","dye-sensitized solar cells","perovskite solar cells","text mining","figure mining","UV/vis absorption"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/69eb0511-090a-4098-8dd1-de9f4e146b39/download","https://www.rioxx.net/licenses/all-rights-reserved/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["10.17863/CAM.65955"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d4b746af-f83d-47ed-8c3f-9e50d838ffa2/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis addresses the use of literature-mining software tools to automatically create databases, to accelerate progress in third-generation photovoltaic devices. New discoveries in this area are typically published in research articles, where the data are contained in figures, tables and text. The field could benefit from a single resource where these key data are stored, such that progress might be driven by new machine-learning techniques, applied to large quantities of structured information. The thesis also describes new tools designed to extract other pertinent data. Chapter 1 discusses the potential of third-generation solar devices as alternative fuel sources. Chapter 2 introduces relevant state-of-the-art data-mining techniques and methods used throughout the thesis. Chapter 3 details a bespoke workflow for the extraction of UV/vis absorption spectral peak data from the literature; specifically the wavelength of maximum absorption, λmax, and the molar extinction coefficient, ε. Application of the workflow results in a database of 18,309 chemical records of experimental data. Chapter 4 describes two databases of photovoltaic properties that were auto-generated from the literature, using a data-extraction pipeline that was tailored for this application. The databases contain over 50,000 records of photovoltaic attributes about dye-sensitized solar cells and perovskite solar cells, and all the key properties were extracted with a collective precision of 84%. Chapter 5 introduces ChemSchematicResolver, a high-throughput data-extraction toolkit that can automatically detect chemical-schematic diagrams inside a document, resolve any R-group substituents and convert the resulting diagram to a machine-readable format. The tool was tested on a new evaluation set and all assessed areas achieved precisions of 83-100%. Chapter 6 discusses contributions to two new independent image-extraction software tools, ImageDataExtractor and FigureDataExtractor. ImageDataExtractor is a toolkit designed to extract quantitative data of particle size, shape and distribution from microscopy images. FigureDataExtractor was built to reproduce quantitative data from line graphs. Chapter 7 concludes the work and investigates potential avenues for future research."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["563e5f5d461110c49c3d1162dbc1ab1b","353adac0d1ebdfd65ab16480263c3c87"]},{"key":"dc:title","label":"Title","values":["Automatic creation of novel databases with data-extraction tools for the advancement of third-generation solar cells."]}]}],"canonical_facts":{"dc:contributor.advisor":["Cole, Jacqueline"],"dc:contributor.sponsor":["Tessella Science and Technology Facilities Council (STFC)"],"dc:creator":["Beard, Edward"],"dc:date.issued":["2020-09-30"],"dc:description.abstract":["This thesis addresses the use of literature-mining software tools to automatically create databases, to accelerate progress in third-generation photovoltaic devices. New discoveries in this area are typically published in research articles, where the data are contained in figures, tables and text. The field could benefit from a single resource where these key data are stored, such that progress might be driven by new machine-learning techniques, applied to large quantities of structured information. The thesis also describes new tools designed to extract other pertinent data. Chapter 1 discusses the potential of third-generation solar devices as alternative fuel sources. Chapter 2 introduces relevant state-of-the-art data-mining techniques and methods used throughout the thesis. Chapter 3 details a bespoke workflow for the extraction of UV/vis absorption spectral peak data from the literature; specifically the wavelength of maximum absorption, λmax, and the molar extinction coefficient, ε. Application of the workflow results in a database of 18,309 chemical records of experimental data. Chapter 4 describes two databases of photovoltaic properties that were auto-generated from the literature, using a data-extraction pipeline that was tailored for this application. The databases contain over 50,000 records of photovoltaic attributes about dye-sensitized solar cells and perovskite solar cells, and all the key properties were extracted with a collective precision of 84%. Chapter 5 introduces ChemSchematicResolver, a high-throughput data-extraction toolkit that can automatically detect chemical-schematic diagrams inside a document, resolve any R-group substituents and convert the resulting diagram to a machine-readable format. The tool was tested on a new evaluation set and all assessed areas achieved precisions of 83-100%. Chapter 6 discusses contributions to two new independent image-extraction software tools, ImageDataExtractor and FigureDataExtractor. ImageDataExtractor is a toolkit designed to extract quantitative data of particle size, shape and distribution from microscopy images. FigureDataExtractor was built to reproduce quantitative data from line graphs. Chapter 7 concludes the work and investigates potential avenues for future research."],"dc:format.checksum.md5":["563e5f5d461110c49c3d1162dbc1ab1b","353adac0d1ebdfd65ab16480263c3c87"],"dc:identifier.doi":["10.17863/CAM.65955"],"dc:identifier.uri":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/d4b746af-f83d-47ed-8c3f-9e50d838ffa2/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/318837"],"dc:rights":["https://apollo8-f-pro.lib.cam.ac.uk/bitstreams/69eb0511-090a-4098-8dd1-de9f4e146b39/download","https://www.rioxx.net/licenses/all-rights-reserved/"],"dc:subject":["data mining","dye-sensitized solar cells","perovskite solar cells","text mining","figure mining","UV/vis absorption"],"dc:title":["Automatic creation of novel databases with data-extraction tools for the advancement of third-generation solar cells."],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:00Z"}