University of Cambridge
Automatic creation of novel databases with data-extraction tools for the advancement of third-generation solar cells.
Abstract
dc:description.abstractThis thesis addresses the use of literature-mining software tools to automatically create databases, to accelerate progress in third-generation photovoltaic devices. New discoveries in this area are typically published in research articles, where the data are contained in figures, tables and text. The field could benefit from a single resource where these key data are stored, such that progress might be driven by new machine-learning techniques, applied to large quantities of structured information. The thesis also describes new tools designed to extract other pertinent data. Chapter 1 discusses the potential of third-generation solar devices as alternative fuel sources. Chapter 2 introduces relevant state-of-the-art data-mining techniques and methods used throughout the thesis. Chapter 3 details a bespoke workflow for the extraction of UV/vis absorption spectral peak data from the literature; specifically the wavelength of maximum absorption, λmax, and the molar extinction coefficient, ε. Application of the workflow results in a database of 18,309 chemical records of experimental data. Chapter 4 describes two databases of photovoltaic properties that were auto-generated from the literature, using a data-extraction pipeline that was tailored for this application. The databases contain over 50,000 records of photovoltaic attributes about dye-sensitized solar cells and perovskite solar cells, and all the key properties were extracted with a collective precision of 84%. Chapter 5 introduces ChemSchematicResolver, a high-throughput data-extraction toolkit that can automatically detect chemical-schematic diagrams inside a document, resolve any R-group substituents and convert the resulting diagram to a machine-readable format. The tool was tested on a new evaluation set and all assessed areas achieved precisions of 83-100%. Chapter 6 discusses contributions to two new independent image-extraction software tools, ImageDataExtractor and FigureDataExtractor. ImageDataExtractor is a toolkit designed to extract quantitative data of particle size, shape and distribution from microscopy images. FigureDataExtractor was built to reproduce quantitative data from line graphs. Chapter 7 concludes the work and investigates potential avenues for future research.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2020
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Beard, Edward
- Advisor dc:contributor.advisor
-
- Cole, Jacqueline
Subjects
dc:subject × 6Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.65955
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/318837