Back to results

University of Cambridge

Extended many-facet Rasch models: Accounting for rater effects in automated essay scoring systems

Abstract

dc:description.abstract

In this thesis, we investigate the problem of unreliable labels in supervised machine learning in the context of automated essay scoring (AES) systems. Specifically, we investigate approaches to identifying and accounting for systemic error by human raters which can lead to the training of AES systems on biased training labels, resulting in bias in the trained AES system. Previous research has developed many-facet Rasch models (MFRMs) that can model and account for rater severity; however, systemic rater behaviours are often non-uniform across scales and assessment criteria. I argue that models that can capture such behaviours are indicated in order to account for non-uniform rater behaviour; I also argue that these models provide ratings that are more consistent with principles of measurement theory, independent of raters. This thesis comprises the following contributions: Novel models for capturing rater severity and novel estimation algorithm: • Novel extended MFRM forms capable of identifying and accounting for a range of non-uniform rater behaviours. • A novel Rasch estimation algorithm which builds on previous non-iterative conditional approaches to Rasch parameter estimation: the conditional pairwise adjacent thresholds (CPAT ) algorithm. Evaluation of the novel models: • An evaluation of the efficiency and efficacy of CPAT and extended MFRMs using simulated data sets. • A case study in which extended MFRMs are applied to a real data set that has previously been analysed in the literature, demonstrating the differences in inference obtained from using extended MFRMs. Research code: • RaschPy, a Python package for Rasch analysis developed from the code used for the analysis in this work, with a comprehensive user manual. Using the novel methods, the work answers the following research questions: • Can systemic, non-uniform error in human essay scoring be modelled and accounted for? • What effect does improving the quality of training labels for an automated essay scoring system through modelling and accounting for systemic rater error have on performance? I conclude that extended MFRMs are necessary because: they naturally capture and account for a variety of systemic rater behaviours which the standard MFRM cannot; they are more consistent with principles of measurement; they produce more accurate person estimates; they provide the basis for richer, more nuanced rater feedback, and they help remove a key source of systemic bias in AES system training. Given the import of extended MFRMs, fast accurate methods for deploying them are essential. I demonstrate that CPAT, which I have opensourced through RaschPy, is a fast, accurate estimation algorithm. Taken together, the work I present in this thesis represents a set of powerful tools to enhance rater analysis and AES system training.

Degree

thesis:*
Name dc:type.qualificationname
Doctor of Philosophy (PhD)
Level dc:type.qualificationlevel
Doctoral
Grantor dc:publisher.institution
University of Cambridge
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Elliott, Mark
Advisor dc:contributor.advisor
  • Buttery, Paula J

Subjects

dc:subject × 5

Rights

dc:rights
Language dc:language
eng

Identifiers

dc:identifier.*
Author Identifier
0000-0003-3302-5477
OAI identifier oai:identifier
oai:www.repository.cam.ac.uk:1810/398801

Chain of custody

source
Harvested from
Cambridge University
Base URL
api.repository.cam.ac.uk/server/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Elliott, Mark. Extended many-facet Rasch models: Accounting for rater effects in automated essay scoring systems. Doctoral thesis, University of Cambridge, 2025. https://doi.org/10.17863/CAM.127567