{"id":{"repo_id":"usf","oai_identifier":"oai:digitalcommons.usf.edu:etd-1035"},"canonical_url":"https://search.dev.ndltd.org/etd/usf/oai:digitalcommons.usf.edu:etd-1035","repository":{"repo_id":"usf","name":"University of South Florida","base_url":"https://digitalcommons.usf.edu/do/oai/"},"display":{"title":"Psychometrics of OSCE Standardized Patient Measurements","abstract":"This study examined the reliability and validity of scores taken from a series of four task simulations used to evaluate medical students. The four role-play exercises represented two different cases or scripts, yielding two pairs of exercises that are considered alternate forms. The design allowed examining what is essentially the ceiling for reliability and validity of ratings taken in such role plays. A multitrait-multimethod (MTMM) matrix was computed with exercises as methods and competencies (history taking, clinical skills, and communication) as traits. The results within alternate forms (within cases) were then used as a baseline to evaluate the reliability and validity of scores between the alternate forms (between cases). There was much less of an exercise effect (method variance, monomethod bias) in this study than is typically found in MTMM matrices for performance measurement. However, the convergent validity of the dimensions across exercises was weak both within and between cases. The study also examined the reliability of ratings by training raters to watch video recordings of the same four exercises who then complete the same forms used by the standardized patients. Generalizability analysis was used to compute variance components for case, station, rater, and ratee (medical student), which allowed the computation of reliability estimates for multiple designs. Both the generalizability analysis and the MTMM analysis indicated that rather long examinations (approximately 20 to 40 exercises) would be needed to create reliable examination scores for this population of examinees. Additionally, interjudge agreement was better for more objective dimensions (history taking, physical examination) than for the more subjective dimension (communication).","abstract_html":"This study examined the reliability and validity of scores taken from a series of four task simulations used to evaluate medical students. The four role-play exercises represented two different cases or scripts, yielding two pairs of exercises that are considered alternate forms. The design allowed examining what is essentially the ceiling for reliability and validity of ratings taken in such role plays. A multitrait-multimethod (MTMM) matrix was computed with exercises as methods and competencies (history taking, clinical skills, and communication) as traits. The results within alternate forms (within cases) were then used as a baseline to evaluate the reliability and validity of scores between the alternate forms (between cases). There was much less of an exercise effect (method variance, monomethod bias) in this study than is typically found in MTMM matrices for performance measurement. However, the convergent validity of the dimensions across exercises was weak both within and between cases. The study also examined the reliability of ratings by training raters to watch video recordings of the same four exercises who then complete the same forms used by the standardized patients. Generalizability analysis was used to compute variance components for case, station, rater, and ratee (medical student), which allowed the computation of reliability estimates for multiple designs. Both the generalizability analysis and the MTMM analysis indicated that rather long examinations (approximately 20 to 40 exercises) would be needed to create reliable examination scores for this population of examinees. Additionally, interjudge agreement was better for more objective dimensions (history taking, physical examination) than for the more subjective dimension (communication).","abstract_has_math":false,"creators":["Stilson, Frederick R. B"],"institution":"Digital Commons @ University of South Florida","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2008,"date_issued":"2008-05-09T07:00:00Z","date_published":"2008-05-09T07:00:00Z","updated_at":"2026-07-24T05:42:11Z","subjects":["Reliability","Validity","Assessment centers","Interdisciplinary","Medicine","American Studies","Arts and Humanities"],"languages":[],"rights":["default"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.usf.edu/etd/36","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Stilson, Frederick R. B"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2008-05-09T07:00:00Z"]},{"key":"dc:publisher","label":"Institution","values":["Digital Commons @ University of South Florida"]},{"key":"dc:type","label":"Dc Type","values":["dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Reliability","Validity","Assessment centers","Interdisciplinary","Medicine","American Studies","Arts and Humanities"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["default"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.usf.edu/etd/36","https://digitalcommons.usf.edu/context/etd/article/1035/viewcontent/etd__36.pdf"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["This study examined the reliability and validity of scores taken from a series of four task simulations used to evaluate medical students. The four role-play exercises represented two different cases or scripts, yielding two pairs of exercises that are considered alternate forms. The design allowed examining what is essentially the ceiling for reliability and validity of ratings taken in such role plays. A multitrait-multimethod (MTMM) matrix was computed with exercises as methods and competencies (history taking, clinical skills, and communication) as traits. The results within alternate forms (within cases) were then used as a baseline to evaluate the reliability and validity of scores between the alternate forms (between cases). There was much less of an exercise effect (method variance, monomethod bias) in this study than is typically found in MTMM matrices for performance measurement. However, the convergent validity of the dimensions across exercises was weak both within and between cases. The study also examined the reliability of ratings by training raters to watch video recordings of the same four exercises who then complete the same forms used by the standardized patients. Generalizability analysis was used to compute variance components for case, station, rater, and ratee (medical student), which allowed the computation of reliability estimates for multiple designs. Both the generalizability analysis and the MTMM analysis indicated that rather long examinations (approximately 20 to 40 exercises) would be needed to create reliable examination scores for this population of examinees. Additionally, interjudge agreement was better for more objective dimensions (history taking, physical examination) than for the more subjective dimension (communication)."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:source","label":"Dc Source","values":["USF Tampa Graduate Theses and Dissertations"]},{"key":"dc:title","label":"Title","values":["Psychometrics of OSCE Standardized Patient Measurements"]}]}],"canonical_facts":{"dc:creator":["Stilson, Frederick R. B"],"dc:date":["2008-05-09T07:00:00Z"],"dc:description":["This study examined the reliability and validity of scores taken from a series of four task simulations used to evaluate medical students. The four role-play exercises represented two different cases or scripts, yielding two pairs of exercises that are considered alternate forms. The design allowed examining what is essentially the ceiling for reliability and validity of ratings taken in such role plays. A multitrait-multimethod (MTMM) matrix was computed with exercises as methods and competencies (history taking, clinical skills, and communication) as traits. The results within alternate forms (within cases) were then used as a baseline to evaluate the reliability and validity of scores between the alternate forms (between cases). There was much less of an exercise effect (method variance, monomethod bias) in this study than is typically found in MTMM matrices for performance measurement. However, the convergent validity of the dimensions across exercises was weak both within and between cases. The study also examined the reliability of ratings by training raters to watch video recordings of the same four exercises who then complete the same forms used by the standardized patients. Generalizability analysis was used to compute variance components for case, station, rater, and ratee (medical student), which allowed the computation of reliability estimates for multiple designs. Both the generalizability analysis and the MTMM analysis indicated that rather long examinations (approximately 20 to 40 exercises) would be needed to create reliable examination scores for this population of examinees. Additionally, interjudge agreement was better for more objective dimensions (history taking, physical examination) than for the more subjective dimension (communication)."],"dc:format":["application/pdf"],"dc:identifier":["https://digitalcommons.usf.edu/etd/36","https://digitalcommons.usf.edu/context/etd/article/1035/viewcontent/etd__36.pdf"],"dc:publisher":["Digital Commons @ University of South Florida"],"dc:rights":["default"],"dc:source":["USF Tampa Graduate Theses and Dissertations"],"dc:subject":["Reliability","Validity","Assessment centers","Interdisciplinary","Medicine","American Studies","Arts and Humanities"],"dc:title":["Psychometrics of OSCE Standardized Patient Measurements"],"dc:type":["dissertation"]},"updated_at":"2026-07-24T05:42:11Z"}