Back to results

University of Cambridge

Automatic creation of novel databases with data-extraction tools for the advancement of third-generation solar cells.

Abstract

dc:description.abstract

This thesis addresses the use of literature-mining software tools to automatically create databases, to accelerate progress in third-generation photovoltaic devices. New discoveries in this area are typically published in research articles, where the data are contained in figures, tables and text. The field could benefit from a single resource where these key data are stored, such that progress might be driven by new machine-learning techniques, applied to large quantities of structured information. The thesis also describes new tools designed to extract other pertinent data. Chapter 1 discusses the potential of third-generation solar devices as alternative fuel sources. Chapter 2 introduces relevant state-of-the-art data-mining techniques and methods used throughout the thesis. Chapter 3 details a bespoke workflow for the extraction of UV/vis absorption spectral peak data from the literature; specifically the wavelength of maximum absorption, λmax, and the molar extinction coefficient, ε. Application of the workflow results in a database of 18,309 chemical records of experimental data. Chapter 4 describes two databases of photovoltaic properties that were auto-generated from the literature, using a data-extraction pipeline that was tailored for this application. The databases contain over 50,000 records of photovoltaic attributes about dye-sensitized solar cells and perovskite solar cells, and all the key properties were extracted with a collective precision of 84%. Chapter 5 introduces ChemSchematicResolver, a high-throughput data-extraction toolkit that can automatically detect chemical-schematic diagrams inside a document, resolve any R-group substituents and convert the resulting diagram to a machine-readable format. The tool was tested on a new evaluation set and all assessed areas achieved precisions of 83-100%. Chapter 6 discusses contributions to two new independent image-extraction software tools, ImageDataExtractor and FigureDataExtractor. ImageDataExtractor is a toolkit designed to extract quantitative data of particle size, shape and distribution from microscopy images. FigureDataExtractor was built to reproduce quantitative data from line graphs. Chapter 7 concludes the work and investigates potential avenues for future research.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2020

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Beard, Edward
Advisor dc:contributor.advisor
  • Cole, Jacqueline

Subjects

dc:subject × 6

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.65955
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/318837

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Beard, Edward. Automatic creation of novel databases with data-extraction tools for the advancement of third-generation solar cells.. Doctoral thesis, University of Cambridge, 2020. https://doi.org/10.17863/CAM.65955