{"id":{"repo_id":"cambridge","oai_identifier":"oai:www.repository.cam.ac.uk:1810/398801"},"canonical_url":"https://search.dev.ndltd.org/etd/cambridge/oai:www.repository.cam.ac.uk:1810/398801","repository":{"repo_id":"cambridge","name":"Cambridge University","base_url":"https://api.repository.cam.ac.uk/server/oai/request"},"display":{"title":"Extended many-facet Rasch models: Accounting for rater effects in automated essay scoring systems","abstract":"In this thesis, we investigate the problem of unreliable labels in supervised machine learning in the context of automated essay scoring (AES) systems. Specifically, we investigate approaches to identifying and accounting for systemic error by human raters which can lead to the training of AES systems on biased training labels, resulting in bias in the trained AES system. Previous research has developed many-facet Rasch models (MFRMs) that can model and account for rater severity; however, systemic rater behaviours are often non-uniform across scales and assessment criteria. I argue that models that can capture such behaviours are indicated in order to account for non-uniform rater behaviour; I also argue that these models provide ratings that are more consistent with principles of measurement theory, independent of raters. This thesis comprises the following contributions: Novel models for capturing rater severity and novel estimation algorithm: • Novel extended MFRM forms capable of identifying and accounting for a range of non-uniform rater behaviours. • A novel Rasch estimation algorithm which builds on previous non-iterative conditional approaches to Rasch parameter estimation: the conditional pairwise adjacent thresholds (CPAT ) algorithm. Evaluation of the novel models: • An evaluation of the efficiency and efficacy of CPAT and extended MFRMs using simulated data sets. • A case study in which extended MFRMs are applied to a real data set that has previously been analysed in the literature, demonstrating the differences in inference obtained from using extended MFRMs. Research code: • RaschPy, a Python package for Rasch analysis developed from the code used for the analysis in this work, with a comprehensive user manual. Using the novel methods, the work answers the following research questions: • Can systemic, non-uniform error in human essay scoring be modelled and accounted for? • What effect does improving the quality of training labels for an automated essay scoring system through modelling and accounting for systemic rater error have on performance? I conclude that extended MFRMs are necessary because: they naturally capture and account for a variety of systemic rater behaviours which the standard MFRM cannot; they are more consistent with principles of measurement; they produce more accurate person estimates; they provide the basis for richer, more nuanced rater feedback, and they help remove a key source of systemic bias in AES system training. Given the import of extended MFRMs, fast accurate methods for deploying them are essential. I demonstrate that CPAT, which I have opensourced through RaschPy, is a fast, accurate estimation algorithm. Taken together, the work I present in this thesis represents a set of powerful tools to enhance rater analysis and AES system training.","abstract_html":"In this thesis, we investigate the problem of unreliable labels in supervised machine learning in the context of automated essay scoring (AES) systems. Specifically, we investigate approaches to identifying and accounting for systemic error by human raters which can lead to the training of AES systems on biased training labels, resulting in bias in the trained AES system. Previous research has developed many-facet Rasch models (MFRMs) that can model and account for rater severity; however, systemic rater behaviours are often non-uniform across scales and assessment criteria. I argue that models that can capture such behaviours are indicated in order to account for non-uniform rater behaviour; I also argue that these models provide ratings that are more consistent with principles of measurement theory, independent of raters. This thesis comprises the following contributions: Novel models for capturing rater severity and novel estimation algorithm: • Novel extended MFRM forms capable of identifying and accounting for a range of non-uniform rater behaviours. • A novel Rasch estimation algorithm which builds on previous non-iterative conditional approaches to Rasch parameter estimation: the conditional pairwise adjacent thresholds (CPAT ) algorithm. Evaluation of the novel models: • An evaluation of the efficiency and efficacy of CPAT and extended MFRMs using simulated data sets. • A case study in which extended MFRMs are applied to a real data set that has previously been analysed in the literature, demonstrating the differences in inference obtained from using extended MFRMs. Research code: • RaschPy, a Python package for Rasch analysis developed from the code used for the analysis in this work, with a comprehensive user manual. Using the novel methods, the work answers the following research questions: • Can systemic, non-uniform error in human essay scoring be modelled and accounted for? • What effect does improving the quality of training labels for an automated essay scoring system through modelling and accounting for systemic rater error have on performance? I conclude that extended MFRMs are necessary because: they naturally capture and account for a variety of systemic rater behaviours which the standard MFRM cannot; they are more consistent with principles of measurement; they produce more accurate person estimates; they provide the basis for richer, more nuanced rater feedback, and they help remove a key source of systemic bias in AES system training. Given the import of extended MFRMs, fast accurate methods for deploying them are essential. I demonstrate that CPAT, which I have opensourced through RaschPy, is a fast, accurate estimation algorithm. Taken together, the work I present in this thesis represents a set of powerful tools to enhance rater analysis and AES system training.","abstract_has_math":false,"creators":["Elliott, Mark"],"institution":"University of Cambridge","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Buttery, Paula J"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-03-12","date_published":"2025-03-12","updated_at":"2026-07-22T22:24:27Z","subjects":["Automated essay scoring","Measurement theory","Parameter estimation algorithm","Rasch","Rater behaviour"],"languages":["eng"],"rights":[],"rights_urls":["https://www.repository.cam.ac.uk/bitstreams/282c9d15-fdf6-4a4f-bdf4-778bef6b537a/download","http://purl.org/NET/rdflicense/allrightsreserved"],"identifier_entries":[{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000333025477"],"render_values":[{"text":"0000-0003-3302-5477","href":"https://orcid.org/0000-0003-3302-5477","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.17863/CAM.127567","outbound_label":"DOI","outbound_source":"dc:identifier.doi"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Buttery, Paula J"]},{"key":"dc:contributor.sponsor","label":"Sponsor","values":["Course fees paid by employers: Cambridge University Press & Assessment"]},{"key":"dc:creator","label":"Author","values":["Elliott, Mark"]},{"key":"dc:creator.authoridentifier","label":"Author Identifier","values":["0000000333025477"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2025-03-12"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["University of Cambridge"]},{"key":"dc:relation.isreferencedby.uri","label":"Dc Relation Isreferencedby URI","values":["https://www.repository.cam.ac.uk/handle/1810/398801"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Automated essay scoring","Measurement theory","Parameter estimation algorithm","Rasch","Rater behaviour"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["https://www.repository.cam.ac.uk/bitstreams/282c9d15-fdf6-4a4f-bdf4-778bef6b537a/download","http://purl.org/NET/rdflicense/allrightsreserved"]},{"key":"dc:rights.embargodate","label":"Dc Rights Embargodate","values":["2027-02-27"]},{"key":"dc:rights.embargotype","label":"Dc Rights Embargotype","values":["embargo"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.doi","label":"DOI","values":["https://doi.org/10.17863/CAM.127567"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://www.repository.cam.ac.uk/bitstreams/f3f910b1-6808-4c3b-bb57-9576bfe216ab/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis, we investigate the problem of unreliable labels in supervised machine learning in the context of automated essay scoring (AES) systems. Specifically, we investigate approaches to identifying and accounting for systemic error by human raters which can lead to the training of AES systems on biased training labels, resulting in bias in the trained AES system. Previous research has developed many-facet Rasch models (MFRMs) that can model and account for rater severity; however, systemic rater behaviours are often non-uniform across scales and assessment criteria. I argue that models that can capture such behaviours are indicated in order to account for non-uniform rater behaviour; I also argue that these models provide ratings that are more consistent with principles of measurement theory, independent of raters. This thesis comprises the following contributions: Novel models for capturing rater severity and novel estimation algorithm: • Novel extended MFRM forms capable of identifying and accounting for a range of non-uniform rater behaviours. • A novel Rasch estimation algorithm which builds on previous non-iterative conditional approaches to Rasch parameter estimation: the conditional pairwise adjacent thresholds (CPAT ) algorithm. Evaluation of the novel models: • An evaluation of the efficiency and efficacy of CPAT and extended MFRMs using simulated data sets. • A case study in which extended MFRMs are applied to a real data set that has previously been analysed in the literature, demonstrating the differences in inference obtained from using extended MFRMs. Research code: • RaschPy, a Python package for Rasch analysis developed from the code used for the analysis in this work, with a comprehensive user manual. Using the novel methods, the work answers the following research questions: • Can systemic, non-uniform error in human essay scoring be modelled and accounted for? • What effect does improving the quality of training labels for an automated essay scoring system through modelling and accounting for systemic rater error have on performance? I conclude that extended MFRMs are necessary because: they naturally capture and account for a variety of systemic rater behaviours which the standard MFRM cannot; they are more consistent with principles of measurement; they produce more accurate person estimates; they provide the basis for richer, more nuanced rater feedback, and they help remove a key source of systemic bias in AES system training. Given the import of extended MFRMs, fast accurate methods for deploying them are essential. I demonstrate that CPAT, which I have opensourced through RaschPy, is a fast, accurate estimation algorithm. Taken together, the work I present in this thesis represents a set of powerful tools to enhance rater analysis and AES system training."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["45b332a8cbed7da0f9d5d08d92f814b1","87eda9de84448d1f82354d60eee3eb5f"]},{"key":"dc:title","label":"Title","values":["Extended many-facet Rasch models: Accounting for rater effects in automated essay scoring systems"]}]}],"canonical_facts":{"dc:contributor.advisor":["Buttery, Paula J"],"dc:contributor.sponsor":["Course fees paid by employers: Cambridge University Press & Assessment"],"dc:creator":["Elliott, Mark"],"dc:creator.authoridentifier":["0000000333025477"],"dc:date.issued":["2025-03-12"],"dc:description.abstract":["In this thesis, we investigate the problem of unreliable labels in supervised machine learning in the context of automated essay scoring (AES) systems. Specifically, we investigate approaches to identifying and accounting for systemic error by human raters which can lead to the training of AES systems on biased training labels, resulting in bias in the trained AES system. Previous research has developed many-facet Rasch models (MFRMs) that can model and account for rater severity; however, systemic rater behaviours are often non-uniform across scales and assessment criteria. I argue that models that can capture such behaviours are indicated in order to account for non-uniform rater behaviour; I also argue that these models provide ratings that are more consistent with principles of measurement theory, independent of raters. This thesis comprises the following contributions: Novel models for capturing rater severity and novel estimation algorithm: • Novel extended MFRM forms capable of identifying and accounting for a range of non-uniform rater behaviours. • A novel Rasch estimation algorithm which builds on previous non-iterative conditional approaches to Rasch parameter estimation: the conditional pairwise adjacent thresholds (CPAT ) algorithm. Evaluation of the novel models: • An evaluation of the efficiency and efficacy of CPAT and extended MFRMs using simulated data sets. • A case study in which extended MFRMs are applied to a real data set that has previously been analysed in the literature, demonstrating the differences in inference obtained from using extended MFRMs. Research code: • RaschPy, a Python package for Rasch analysis developed from the code used for the analysis in this work, with a comprehensive user manual. Using the novel methods, the work answers the following research questions: • Can systemic, non-uniform error in human essay scoring be modelled and accounted for? • What effect does improving the quality of training labels for an automated essay scoring system through modelling and accounting for systemic rater error have on performance? I conclude that extended MFRMs are necessary because: they naturally capture and account for a variety of systemic rater behaviours which the standard MFRM cannot; they are more consistent with principles of measurement; they produce more accurate person estimates; they provide the basis for richer, more nuanced rater feedback, and they help remove a key source of systemic bias in AES system training. Given the import of extended MFRMs, fast accurate methods for deploying them are essential. I demonstrate that CPAT, which I have opensourced through RaschPy, is a fast, accurate estimation algorithm. Taken together, the work I present in this thesis represents a set of powerful tools to enhance rater analysis and AES system training."],"dc:format.checksum.md5":["45b332a8cbed7da0f9d5d08d92f814b1","87eda9de84448d1f82354d60eee3eb5f"],"dc:identifier.doi":["https://doi.org/10.17863/CAM.127567"],"dc:identifier.uri":["https://www.repository.cam.ac.uk/bitstreams/f3f910b1-6808-4c3b-bb57-9576bfe216ab/download"],"dc:language":["eng"],"dc:publisher.institution":["University of Cambridge"],"dc:relation.isreferencedby.uri":["https://www.repository.cam.ac.uk/handle/1810/398801"],"dc:rights":["https://www.repository.cam.ac.uk/bitstreams/282c9d15-fdf6-4a4f-bdf4-778bef6b537a/download","http://purl.org/NET/rdflicense/allrightsreserved"],"dc:rights.embargodate":["2027-02-27"],"dc:rights.embargotype":["embargo"],"dc:subject":["Automated essay scoring","Measurement theory","Parameter estimation algorithm","Rasch","Rater behaviour"],"dc:title":["Extended many-facet Rasch models: Accounting for rater effects in automated essay scoring systems"],"dc:type":["Thesis"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-22T22:24:27Z"}