{"id":{"repo_id":"south-carolina","oai_identifier":"oai:scholarcommons.sc.edu:etd-3636"},"canonical_url":"https://search.dev.ndltd.org/etd/south-carolina/oai:scholarcommons.sc.edu:etd-3636","repository":{"repo_id":"south-carolina","name":"University of South Carolina","base_url":"https://scholarcommons.sc.edu/do/oai/"},"display":{"title":"Model-Based Measures of Interrater Agreement","abstract":"<p> Cohen's &kappa; (1960) is almost universally used for the assessment of the strength of agreement among raters, each classifying subjects as e.g. diseased or not. Nelson and Edwards (2008) propose a generalized linear mixed model for the agreement process, showing that Cohen's &kappa; can seriously underestimate agreement and proposing a model-based coefficient &kappa;<sub>M</sub> which does not suffer from this flaw (at least when the assumed model is correct). They discuss computationally intensive methods for estimation and inference on &kappa;<sub>M</sub>. This paper builds on this previous work, adding theoretical justification and practical methods for estimation and inference. Their model for the agreement process is motivated by a threshold model with latent crossed item - rater random effects plus interaction and pure error. Under this framework it is shown that &kappa;<sub>M</sub> is a monotone function of the Pearson (1900; 1913) tetrachoric correlation &rho; between the latent effects of any two raters. A practical method for estimation and inference on &kappa;<sub>M</sub> which bypasses computationally- intensive model fitting is defined and studied; this approach is analogous to that proposed by Pearson for estimation of &rho; in a 2 × 2 table. Asymptotic estimator behavior is established using a generalization of Lehmann's theory of two-sample U-statistics (1951). Standard errors are defined using this asymptotic approximation, Owen's pigeonhole bootstrap (2007) and a counting statistics approach. Simulation studies suggest that the new estimation methods have negligible bias and with the bootstrap approach can perform well in terms of confidence interval coverage when the number of items/raters is on the order of 60 or more. For the robustness of the new estimation method, the bootstrap approach can perform well under the assumption that both random effects follow a t distribution with moderate-to-large degree of freedom (df > 20), with the probit link function. Under the logit link function, or a probit link with low degrees of freedom (df = 3) for either of the random effects, actual coverage of nominal 90% confidence intervals for &rho; could be unacceptable low.</p> <p>A function KappaM written in R (R Development Core Team, 2005) is provided which creates agreement image plots and calculates parameter estimates with bootstrap standard errors. The methods are illustrated with both a prostate biopsy example (Allsbrook et al., 2001) and a mammography example (Beam et al., 1996, 2003); the latter example shows substantial differences between &kappa;<sub>M</sub> and &kappa; under high prevalence. </p>","abstract_html":"&lt;p&gt; Cohen&#x27;s &amp;kappa; (1960) is almost universally used for the assessment of the strength of agreement among raters, each classifying subjects as e.g. diseased or not. Nelson and Edwards (2008) propose a generalized linear mixed model for the agreement process, showing that Cohen&#x27;s &amp;kappa; can seriously underestimate agreement and proposing a model-based coefficient &amp;kappa;&lt;sub&gt;M&lt;/sub&gt; which does not suffer from this flaw (at least when the assumed model is correct). They discuss computationally intensive methods for estimation and inference on &amp;kappa;&lt;sub&gt;M&lt;/sub&gt;. This paper builds on this previous work, adding theoretical justification and practical methods for estimation and inference. Their model for the agreement process is motivated by a threshold model with latent crossed item - rater random effects plus interaction and pure error. Under this framework it is shown that &amp;kappa;&lt;sub&gt;M&lt;/sub&gt; is a monotone function of the Pearson (1900; 1913) tetrachoric correlation &amp;rho; between the latent effects of any two raters. A practical method for estimation and inference on &amp;kappa;&lt;sub&gt;M&lt;/sub&gt; which bypasses computationally- intensive model fitting is defined and studied; this approach is analogous to that proposed by Pearson for estimation of &amp;rho; in a 2 × 2 table. Asymptotic estimator behavior is established using a generalization of Lehmann&#x27;s theory of two-sample U-statistics (1951). Standard errors are defined using this asymptotic approximation, Owen&#x27;s pigeonhole bootstrap (2007) and a counting statistics approach. Simulation studies suggest that the new estimation methods have negligible bias and with the bootstrap approach can perform well in terms of confidence interval coverage when the number of items/raters is on the order of 60 or more. For the robustness of the new estimation method, the bootstrap approach can perform well under the assumption that both random effects follow a t distribution with moderate-to-large degree of freedom (df &gt; 20), with the probit link function. Under the logit link function, or a probit link with low degrees of freedom (df = 3) for either of the random effects, actual coverage of nominal 90% confidence intervals for &amp;rho; could be unacceptable low.&lt;/p&gt; &lt;p&gt;A function KappaM written in R (R Development Core Team, 2005) is provided which creates agreement image plots and calculates parameter estimates with bootstrap standard errors. The methods are illustrated with both a prostate biopsy example (Allsbrook et al., 2001) and a mammography example (Beam et al., 1996, 2003); the latter example shows substantial differences between &amp;kappa;&lt;sub&gt;M&lt;/sub&gt; and &amp;kappa; under high prevalence. &lt;/p&gt;","abstract_has_math":false,"creators":["Gao, Jie"],"institution":null,"degree_name":"Ph.D.","degree_level":"Campus Access Dissertation","degree_discipline":"Statistics","degree_department":null,"school":null,"contributors":["Don Edwards"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2012,"date_issued":"2012-01-01T08:00:00Z","date_published":"2012-01-01T08:00:00Z","updated_at":"2026-07-24T04:38:23Z","subjects":["Physical Sciences and Mathematics","Statistics and Probability","Statistics","Model-Based Measures","Interrater Agreement"],"languages":[],"rights":["© 2012, Jie Gao"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://scholarcommons.sc.edu/etd/2605","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Don Edwards"]},{"key":"dc:creator","label":"Author","values":["Gao, Jie"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["1970-01-01T08:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Campus Access Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Physical Sciences and Mathematics","Statistics and Probability","Statistics","Model-Based Measures","Interrater Agreement"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["© 2012, Jie Gao"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://scholarcommons.sc.edu/etd/2605"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p> Cohen's &kappa; (1960) is almost universally used for the assessment of the strength of agreement among raters, each classifying subjects as e.g. diseased or not. Nelson and Edwards (2008) propose a generalized linear mixed model for the agreement process, showing that Cohen's &kappa; can seriously underestimate agreement and proposing a model-based coefficient &kappa;<sub>M</sub> which does not suffer from this flaw (at least when the assumed model is correct). They discuss computationally intensive methods for estimation and inference on &kappa;<sub>M</sub>. This paper builds on this previous work, adding theoretical justification and practical methods for estimation and inference. Their model for the agreement process is motivated by a threshold model with latent crossed item - rater random effects plus interaction and pure error. Under this framework it is shown that &kappa;<sub>M</sub> is a monotone function of the Pearson (1900; 1913) tetrachoric correlation &rho; between the latent effects of any two raters. A practical method for estimation and inference on &kappa;<sub>M</sub> which bypasses computationally- intensive model fitting is defined and studied; this approach is analogous to that proposed by Pearson for estimation of &rho; in a 2 × 2 table. Asymptotic estimator behavior is established using a generalization of Lehmann's theory of two-sample U-statistics (1951). Standard errors are defined using this asymptotic approximation, Owen's pigeonhole bootstrap (2007) and a counting statistics approach. Simulation studies suggest that the new estimation methods have negligible bias and with the bootstrap approach can perform well in terms of confidence interval coverage when the number of items/raters is on the order of 60 or more. For the robustness of the new estimation method, the bootstrap approach can perform well under the assumption that both random effects follow a t distribution with moderate-to-large degree of freedom (df > 20), with the probit link function. Under the logit link function, or a probit link with low degrees of freedom (df = 3) for either of the random effects, actual coverage of nominal 90% confidence intervals for &rho; could be unacceptable low.</p> <p>A function KappaM written in R (R Development Core Team, 2005) is provided which creates agreement image plots and calculates parameter estimates with bootstrap standard errors. The methods are illustrated with both a prostate biopsy example (Allsbrook et al., 2001) and a mammography example (Beam et al., 1996, 2003); the latter example shows substantial differences between &kappa;<sub>M</sub> and &kappa; under high prevalence. </p>"]},{"key":"dc:title","label":"Title","values":["Model-Based Measures of Interrater Agreement"]}]}],"canonical_facts":{"dc:contributor":["Don Edwards"],"dc:creator":["Gao, Jie"],"dc:date.available":["1970-01-01T08:00:00Z"],"dc:description.abstract":["<p> Cohen's &kappa; (1960) is almost universally used for the assessment of the strength of agreement among raters, each classifying subjects as e.g. diseased or not. Nelson and Edwards (2008) propose a generalized linear mixed model for the agreement process, showing that Cohen's &kappa; can seriously underestimate agreement and proposing a model-based coefficient &kappa;<sub>M</sub> which does not suffer from this flaw (at least when the assumed model is correct). They discuss computationally intensive methods for estimation and inference on &kappa;<sub>M</sub>. This paper builds on this previous work, adding theoretical justification and practical methods for estimation and inference. Their model for the agreement process is motivated by a threshold model with latent crossed item - rater random effects plus interaction and pure error. Under this framework it is shown that &kappa;<sub>M</sub> is a monotone function of the Pearson (1900; 1913) tetrachoric correlation &rho; between the latent effects of any two raters. A practical method for estimation and inference on &kappa;<sub>M</sub> which bypasses computationally- intensive model fitting is defined and studied; this approach is analogous to that proposed by Pearson for estimation of &rho; in a 2 × 2 table. Asymptotic estimator behavior is established using a generalization of Lehmann's theory of two-sample U-statistics (1951). Standard errors are defined using this asymptotic approximation, Owen's pigeonhole bootstrap (2007) and a counting statistics approach. Simulation studies suggest that the new estimation methods have negligible bias and with the bootstrap approach can perform well in terms of confidence interval coverage when the number of items/raters is on the order of 60 or more. For the robustness of the new estimation method, the bootstrap approach can perform well under the assumption that both random effects follow a t distribution with moderate-to-large degree of freedom (df > 20), with the probit link function. Under the logit link function, or a probit link with low degrees of freedom (df = 3) for either of the random effects, actual coverage of nominal 90% confidence intervals for &rho; could be unacceptable low.</p> <p>A function KappaM written in R (R Development Core Team, 2005) is provided which creates agreement image plots and calculates parameter estimates with bootstrap standard errors. The methods are illustrated with both a prostate biopsy example (Allsbrook et al., 2001) and a mammography example (Beam et al., 1996, 2003); the latter example shows substantial differences between &kappa;<sub>M</sub> and &kappa; under high prevalence. </p>"],"dc:identifier":["https://scholarcommons.sc.edu/etd/2605"],"dc:rights":["© 2012, Jie Gao"],"dc:subject":["Physical Sciences and Mathematics","Statistics and Probability","Statistics","Model-Based Measures","Interrater Agreement"],"dc:title":["Model-Based Measures of Interrater Agreement"],"thesis:degree_discipline":["Statistics"],"thesis:degree_level":["Campus Access Dissertation"],"thesis:degree_name":["Ph.D."]},"updated_at":"2026-07-24T04:38:23Z"}