University of Cambridge
Extended many-facet Rasch models: Accounting for rater effects in automated essay scoring systems
Abstract
dc:description.abstractIn this thesis, we investigate the problem of unreliable labels in supervised machine learning in the context of automated essay scoring (AES) systems. Specifically, we investigate approaches to identifying and accounting for systemic error by human raters which can lead to the training of AES systems on biased training labels, resulting in bias in the trained AES system. Previous research has developed many-facet Rasch models (MFRMs) that can model and account for rater severity; however, systemic rater behaviours are often non-uniform across scales and assessment criteria. I argue that models that can capture such behaviours are indicated in order to account for non-uniform rater behaviour; I also argue that these models provide ratings that are more consistent with principles of measurement theory, independent of raters. This thesis comprises the following contributions: Novel models for capturing rater severity and novel estimation algorithm: • Novel extended MFRM forms capable of identifying and accounting for a range of non-uniform rater behaviours. • A novel Rasch estimation algorithm which builds on previous non-iterative conditional approaches to Rasch parameter estimation: the conditional pairwise adjacent thresholds (CPAT ) algorithm. Evaluation of the novel models: • An evaluation of the efficiency and efficacy of CPAT and extended MFRMs using simulated data sets. • A case study in which extended MFRMs are applied to a real data set that has previously been analysed in the literature, demonstrating the differences in inference obtained from using extended MFRMs. Research code: • RaschPy, a Python package for Rasch analysis developed from the code used for the analysis in this work, with a comprehensive user manual. Using the novel methods, the work answers the following research questions: • Can systemic, non-uniform error in human essay scoring be modelled and accounted for? • What effect does improving the quality of training labels for an automated essay scoring system through modelling and accounting for systemic rater error have on performance? I conclude that extended MFRMs are necessary because: they naturally capture and account for a variety of systemic rater behaviours which the standard MFRM cannot; they are more consistent with principles of measurement; they produce more accurate person estimates; they provide the basis for richer, more nuanced rater feedback, and they help remove a key source of systemic bias in AES system training. Given the import of extended MFRMs, fast accurate methods for deploying them are essential. I demonstrate that CPAT, which I have opensourced through RaschPy, is a fast, accurate estimation algorithm. Taken together, the work I present in this thesis represents a set of powerful tools to enhance rater analysis and AES system training.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2025
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Elliott, Mark
- Advisor dc:contributor.advisor
-
- Buttery, Paula J
Subjects
dc:subject × 5Rights
dc:rightsIdentifiers
dc:identifier.*- Author Identifier
- 0000-0003-3302-5477
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/398801