Back to results

University of South Carolina

Model-Based Measures of Interrater Agreement

Abstract

dc:description.abstract

<p> Cohen's &kappa; (1960) is almost universally used for the assessment of the strength of agreement among raters, each classifying subjects as e.g. diseased or not. Nelson and Edwards (2008) propose a generalized linear mixed model for the agreement process, showing that Cohen's &kappa; can seriously underestimate agreement and proposing a model-based coefficient &kappa;<sub>M</sub> which does not suffer from this flaw (at least when the assumed model is correct). They discuss computationally intensive methods for estimation and inference on &kappa;<sub>M</sub>. This paper builds on this previous work, adding theoretical justification and practical methods for estimation and inference. Their model for the agreement process is motivated by a threshold model with latent crossed item - rater random effects plus interaction and pure error. Under this framework it is shown that &kappa;<sub>M</sub> is a monotone function of the Pearson (1900; 1913) tetrachoric correlation &rho; between the latent effects of any two raters. A practical method for estimation and inference on &kappa;<sub>M</sub> which bypasses computationally- intensive model fitting is defined and studied; this approach is analogous to that proposed by Pearson for estimation of &rho; in a 2 × 2 table. Asymptotic estimator behavior is established using a generalization of Lehmann's theory of two-sample U-statistics (1951). Standard errors are defined using this asymptotic approximation, Owen's pigeonhole bootstrap (2007) and a counting statistics approach. Simulation studies suggest that the new estimation methods have negligible bias and with the bootstrap approach can perform well in terms of confidence interval coverage when the number of items/raters is on the order of 60 or more. For the robustness of the new estimation method, the bootstrap approach can perform well under the assumption that both random effects follow a t distribution with moderate-to-large degree of freedom (df > 20), with the probit link function. Under the logit link function, or a probit link with low degrees of freedom (df = 3) for either of the random effects, actual coverage of nominal 90% confidence intervals for &rho; could be unacceptable low.</p> <p>A function KappaM written in R (R Development Core Team, 2005) is provided which creates agreement image plots and calculates parameter estimates with bootstrap standard errors. The methods are illustrated with both a prostate biopsy example (Allsbrook et al., 2001) and a mammography example (Beam et al., 1996, 2003); the latter example shows substantial differences between &kappa;<sub>M</sub> and &kappa; under high prevalence. </p>

Degree

thesis:*
Name thesis:degree_name
Ph.D.
Level thesis:degree_level
Campus Access Dissertation
Discipline thesis:degree_discipline
Statistics
Year dc:date.available
2012

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Gao, Jie
Contributors dc:contributor
  • Don Edwards

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • © 2012, Jie Gao

Identifiers

dc:identifier.*
Repository record dc:identifier
https://scholarcommons.sc.edu/etd/2605
OAI identifier oai:identifier
oai:scholarcommons.sc.edu:etd-3636

Chain of custody

source
Harvested from
University of South Carolina
Base URL
scholarcommons.sc.edu/do/oai/
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Gao, Jie. Model-Based Measures of Interrater Agreement. Campus Access Dissertation thesis, 2012. https://scholarcommons.sc.edu/etd/2605