Back to results

University of Cambridge

Extracting thermoelectric materials information using natural language processing

Abstract

dc:description.abstract

This thesis aims to empower thermoelectric materials discovery through data-driven methods. Materials discovery has been traditionally guided by the intuition of researchers and trial-and-error. This approach can be greatly accelerated by applying data-oriented methods which rely on high-quality domain-specific databases. This thesis presents the information-extraction methods that were developed and applied to create the first automatically generated thermoelectric materials database from the scientific literature. The database was further utilised to algorithmically construct a domain-specific question-answering (QA) materials dataset. In turn, this dataset was used to fine-tune a variety of small language models to support information extraction in the thermoelectric materials domain. Throughout this work, a range of methods and resources have been introduced which are novel in their specialised materials science scope. Chapter 1 describes the current landscape of thermoelectric material research as well as the prospect of their application. It further introduces natural language processing (NLP) methods and their application in materials discovery and information extraction with regards to the target domain. Chapter 2 introduces the information-extraction toolkit, ChemDataExtractor, which was adapted and applied in Chapter 3 of this research to extract a thermoelectric materials database from the scientific literature. It also gives an overview of language models, focusing on BERT, which was utilised in Chapter 4 and Chapter 5. The following four chapters pertain to the main work and results of this thesis: Chapter 3 focuses on the process of auto-generating a thermoelectric materials database from the scientific literature, using ChemDataExtractor, as well as the insights and analysis that have been gained thanks to this database. Chapter 4 is centred around utilising the previously generated database to algorithmically create a large domain-specific QA dataset. This QA dataset was then used to fine-tune a BERT language model, affording increased performance on a QA task in the thermoelectric materials domain. Furthermore, a method of mixing distinct QA datasets with complementary semantic and syntactic scopes to enhance performance was also introduced. Chapter 5 further improves the performance on the thermoelectrics-specific question answering task by considering different BERT-like models and their ensembles. The ensembling method achieves a significant increase of performance compared to the preceding chapter. Finally, Chapter 6 presents a summary of the methods developed and resources released during this project. It reviews the progress as well as presents avenues for future work and experimentation. Chapter 6 holds the conclusion of the thesis, summarises the results, and outlines potential avenues for future work.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sierepeklis, Odysseas
Advisor dc:contributor.advisor
  • Cole, Jacqueline Manina

Subjects

dc:subject × 4

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
DOI dc:identifier.doi
https://doi.org/10.17863/CAM.121051
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/388912

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Sierepeklis, Odysseas. Extracting thermoelectric materials information using natural language processing. Doctoral thesis, University of Cambridge, 2024. https://doi.org/10.17863/CAM.121051