Back to results

Embry Riddle Aeronautical University

Assessing Reliability of Expert Ratings Among Judges Responding to a Survey Instrument Developed to Study the Long Term Efficacy of the ABET Engineering Criteria, EC2000

Abstract

dc:description.abstract

<p>In today’s assessment processes, especially those evaluations that rely on humans to make subjective judgements, it is necessary to analyze the quality of their ratings. The psychometric issues associated with assessment provide the lens through which researchers interpret results and important decisions are made. Therefore, inter-rater agreement (IRA) and inter-rater reliability (IRR) are pre-requisites for rater-dependent data analysis. A survey instrument cannot provide “good” information if it is not reliable; in other words, reliability is central to the validation of an instrument. When judges cannot be shown to reliably rate a performance, item, or target, the question becomes why the judges’ responses are different from one another. If the judges’ ratings covary unreliably because the construct is poorly defined or the rating framework is defective, then the resultant scores will have questionable meaning. On the other hand, if the judges’ ratings differ because they have a true difference in opinion, this is of importance to the researcher and may not necessarily diminish the validity of the scores. The intraclass correlation coefficient (ICC) is the most efficient method to assess these rater differences and identify the specific sources of inconsistency in measurement. This study examined how ICCs can be used to inform researchers of the extent in which legitimate differences of opinion may appear as a lack of reliability and/or agreement, demonstrating the need for analyzing survey data beyond standard descriptive statistics. Overall, both the IRA and IRR correlations, as calculated by <em>ICC, </em>ranged from .79 to .91 indicating high levels of agreement and consistency in the scoring among the judges' ratings. When group membership was accounted for the IRA values increased suggesting the common judges agreed more than those judges who varied in their perspectives.</p>

Degree

thesis:*
Name thesis:degree_name
Master of Science in Human Factors & Systems
Level thesis:degree_level
Thesis - Open Access
Discipline thesis:degree_discipline
Human Factors and Systems
Year
2006

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Litzinger, Tracy L.
Contributors dc:contributor
  • Shawn Michael Doherty
  • Elizabeth L. Blickensderfer
  • Rosemarie Reynolds

Subjects

dc:subject × 6

Identifiers

dc:identifier.*
Repository record dc:identifier
https://commons.erau.edu/db-theses/123
OAI identifier oai:identifier
oai:commons.erau.edu:db-theses-1170

Chain of custody

source
Harvested from
Embry Riddle Aeronautical University
Base URL
commons.erau.edu/do/oai/
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Litzinger, Tracy L.. Assessing Reliability of Expert Ratings Among Judges Responding to a Survey Instrument Developed to Study the Long Term Efficacy of the ABET Engineering Criteria, EC2000. Thesis - Open Access thesis, 2006. https://commons.erau.edu/db-theses/123