{"id":{"repo_id":"embry-riddle","oai_identifier":"oai:commons.erau.edu:db-theses-1170"},"canonical_url":"https://search.dev.ndltd.org/etd/embry-riddle/oai:commons.erau.edu:db-theses-1170","repository":{"repo_id":"embry-riddle","name":"Embry Riddle Aeronautical University","base_url":"https://commons.erau.edu/do/oai/"},"display":{"title":"Assessing Reliability of Expert Ratings Among Judges Responding to a Survey Instrument Developed to Study the Long Term Efficacy of the ABET Engineering Criteria, EC2000","abstract":"<p>In today’s assessment processes, especially those evaluations that rely on humans to make subjective judgements, it is necessary to analyze the quality of their ratings. The psychometric issues associated with assessment provide the lens through which researchers interpret results and important decisions are made. Therefore, inter-rater agreement (IRA) and inter-rater reliability (IRR) are pre-requisites for rater-dependent data analysis. A survey instrument cannot provide “good” information if it is not reliable; in other words, reliability is central to the validation of an instrument. When judges cannot be shown to reliably rate a performance, item, or target, the question becomes why the judges’ responses are different from one another. If the judges’ ratings covary unreliably because the construct is poorly defined or the rating framework is defective, then the resultant scores will have questionable meaning. On the other hand, if the judges’ ratings differ because they have a true difference in opinion, this is of importance to the researcher and may not necessarily diminish the validity of the scores. The intraclass correlation coefficient (ICC) is the most efficient method to assess these rater differences and identify the specific sources of inconsistency in measurement. This study examined how ICCs can be used to inform researchers of the extent in which legitimate differences of opinion may appear as a lack of reliability and/or agreement, demonstrating the need for analyzing survey data beyond standard descriptive statistics. Overall, both the IRA and IRR correlations, as calculated by <em>ICC, </em>ranged from .79 to .91 indicating high levels of agreement and consistency in the scoring among the judges' ratings. When group membership was accounted for the IRA values increased suggesting the common judges agreed more than those judges who varied in their perspectives.</p>","abstract_html":"&lt;p&gt;In today’s assessment processes, especially those evaluations that rely on humans to make subjective judgements, it is necessary to analyze the quality of their ratings. The psychometric issues associated with assessment provide the lens through which researchers interpret results and important decisions are made. Therefore, inter-rater agreement (IRA) and inter-rater reliability (IRR) are pre-requisites for rater-dependent data analysis. A survey instrument cannot provide “good” information if it is not reliable; in other words, reliability is central to the validation of an instrument. When judges cannot be shown to reliably rate a performance, item, or target, the question becomes why the judges’ responses are different from one another. If the judges’ ratings covary unreliably because the construct is poorly defined or the rating framework is defective, then the resultant scores will have questionable meaning. On the other hand, if the judges’ ratings differ because they have a true difference in opinion, this is of importance to the researcher and may not necessarily diminish the validity of the scores. The intraclass correlation coefficient (ICC) is the most efficient method to assess these rater differences and identify the specific sources of inconsistency in measurement. This study examined how ICCs can be used to inform researchers of the extent in which legitimate differences of opinion may appear as a lack of reliability and/or agreement, demonstrating the need for analyzing survey data beyond standard descriptive statistics. Overall, both the IRA and IRR correlations, as calculated by &lt;em&gt;ICC, &lt;/em&gt;ranged from .79 to .91 indicating high levels of agreement and consistency in the scoring among the judges&#x27; ratings. When group membership was accounted for the IRA values increased suggesting the common judges agreed more than those judges who varied in their perspectives.&lt;/p&gt;","abstract_has_math":false,"creators":["Litzinger, Tracy L."],"institution":null,"degree_name":"Master of Science in Human Factors & Systems","degree_level":"Thesis - Open Access","degree_discipline":"Human Factors and Systems","degree_department":null,"school":null,"contributors":["Shawn Michael Doherty","Elizabeth L. Blickensderfer","Rosemarie Reynolds"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2006,"date_issued":"2006-05-01T07:00:00Z","date_published":"2006-05-01T07:00:00Z","updated_at":"2026-07-27T19:26:28Z","subjects":["expert ratings","ABET","EC2000","long term efficacy","Industrial and Organizational Psychology","Psychology"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://commons.erau.edu/db-theses/123","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Shawn Michael Doherty","Elizabeth L. Blickensderfer","Rosemarie Reynolds"]},{"key":"dc:creator","label":"Author","values":["Litzinger, Tracy L."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"thesis:degree_discipline","label":"Discipline","values":["Human Factors and Systems"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis - Open Access"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science in Human Factors & Systems"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["expert ratings","ABET","EC2000","long term efficacy","Industrial and Organizational Psychology","Psychology"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://commons.erau.edu/db-theses/123"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>In today’s assessment processes, especially those evaluations that rely on humans to make subjective judgements, it is necessary to analyze the quality of their ratings. The psychometric issues associated with assessment provide the lens through which researchers interpret results and important decisions are made. Therefore, inter-rater agreement (IRA) and inter-rater reliability (IRR) are pre-requisites for rater-dependent data analysis. A survey instrument cannot provide “good” information if it is not reliable; in other words, reliability is central to the validation of an instrument. When judges cannot be shown to reliably rate a performance, item, or target, the question becomes why the judges’ responses are different from one another. If the judges’ ratings covary unreliably because the construct is poorly defined or the rating framework is defective, then the resultant scores will have questionable meaning. On the other hand, if the judges’ ratings differ because they have a true difference in opinion, this is of importance to the researcher and may not necessarily diminish the validity of the scores. The intraclass correlation coefficient (ICC) is the most efficient method to assess these rater differences and identify the specific sources of inconsistency in measurement. This study examined how ICCs can be used to inform researchers of the extent in which legitimate differences of opinion may appear as a lack of reliability and/or agreement, demonstrating the need for analyzing survey data beyond standard descriptive statistics. Overall, both the IRA and IRR correlations, as calculated by <em>ICC, </em>ranged from .79 to .91 indicating high levels of agreement and consistency in the scoring among the judges' ratings. When group membership was accounted for the IRA values increased suggesting the common judges agreed more than those judges who varied in their perspectives.</p>"]},{"key":"dc:title","label":"Title","values":["Assessing Reliability of Expert Ratings Among Judges Responding to a Survey Instrument Developed to Study the Long Term Efficacy of the ABET Engineering Criteria, EC2000"]}]}],"canonical_facts":{"dc:contributor":["Shawn Michael Doherty","Elizabeth L. Blickensderfer","Rosemarie Reynolds"],"dc:creator":["Litzinger, Tracy L."],"dc:description.abstract":["<p>In today’s assessment processes, especially those evaluations that rely on humans to make subjective judgements, it is necessary to analyze the quality of their ratings. The psychometric issues associated with assessment provide the lens through which researchers interpret results and important decisions are made. Therefore, inter-rater agreement (IRA) and inter-rater reliability (IRR) are pre-requisites for rater-dependent data analysis. A survey instrument cannot provide “good” information if it is not reliable; in other words, reliability is central to the validation of an instrument. When judges cannot be shown to reliably rate a performance, item, or target, the question becomes why the judges’ responses are different from one another. If the judges’ ratings covary unreliably because the construct is poorly defined or the rating framework is defective, then the resultant scores will have questionable meaning. On the other hand, if the judges’ ratings differ because they have a true difference in opinion, this is of importance to the researcher and may not necessarily diminish the validity of the scores. The intraclass correlation coefficient (ICC) is the most efficient method to assess these rater differences and identify the specific sources of inconsistency in measurement. This study examined how ICCs can be used to inform researchers of the extent in which legitimate differences of opinion may appear as a lack of reliability and/or agreement, demonstrating the need for analyzing survey data beyond standard descriptive statistics. Overall, both the IRA and IRR correlations, as calculated by <em>ICC, </em>ranged from .79 to .91 indicating high levels of agreement and consistency in the scoring among the judges' ratings. When group membership was accounted for the IRA values increased suggesting the common judges agreed more than those judges who varied in their perspectives.</p>"],"dc:identifier":["https://commons.erau.edu/db-theses/123"],"dc:subject":["expert ratings","ABET","EC2000","long term efficacy","Industrial and Organizational Psychology","Psychology"],"dc:title":["Assessing Reliability of Expert Ratings Among Judges Responding to a Survey Instrument Developed to Study the Long Term Efficacy of the ABET Engineering Criteria, EC2000"],"thesis:degree_discipline":["Human Factors and Systems"],"thesis:degree_level":["Thesis - Open Access"],"thesis:degree_name":["Master of Science in Human Factors & Systems"]},"updated_at":"2026-07-27T19:26:28Z"}