{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/15599"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/15599","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The development and evaluation of a systematic training program for increasing both rater reliability and rating accuracy","abstract":"The primary purposes of this study are to identify the characteristics of modeling a rater training program and to develop an efficient training model at the University of Illinois at Urbana-Champaign. This study focuses substantially on a basic conception of rater reliability including true score measurements of examinees’ language proficiency. This study was conducted based on a definition of rater reliability achieved by the reinterpretation of the various meanings of reliability. For these purposes, a basic framework of standardization was achieved using training theories, and this study proposes that a rater training program can be standardized by accomplishing innovative systematic changes that consider (a) the relevant literature, (b) the test instrument itself, (c) the test procedure, and (d) contextual effects such as the characteristics of the stakeholders, their concerns, the structure of the test, or washback effects of the test use. This study utilized a modified version of Lynch’s program evaluation model (1996; 2003) to collect evidence from different sources, including data drawn from the entire evaluation process ranging from needs analysis to a feedback system based on the final product of the evaluation. The effectiveness of both the training program and the individual performances were identified by incorporating all sources of data collected using measurement theory. Mixed methods were proposed for the data analysis. The data analysis involved an investigation of training effectiveness by measuring raters’ scoring reliability, and providing a new training program for raters’ professional improvement. Quantitative data analysis was proposed for analyzing the surveys, the rating corpus, and training effectiveness. Qualitative and document analysis were also essential for analyzing relevant training materials and workshop observation as well as exploring the degree of change in the perceptions of the raters. The results of this study provide educational implications for language testing. At the program level, standardized training contributed to shared responsibilities among test users. The results support the idea that the professionalism of the raters could be improved by providing access to similar quality input which can reinforce their learning and skills via training. The salient value of this dissertation is the collaboration with stakeholders in a test administration situation. Stakeholders’ concerns and challenges were clearly identified, shared, and resolved with the practitioners (the EPT trainer and raters). In addition, I recognize the importance of a balance between understanding fundamental theoretical underpinnings and applying theory through practical experience. It could be concluded that this study contributes to the enhancement of rating validity and the cumulative growth in scoring reliability, as well as a positive washback effect for the future rater training program.","abstract_html":"The primary purposes of this study are to identify the characteristics of modeling a rater training program and to develop an efficient training model at the University of Illinois at Urbana-Champaign. This study focuses substantially on a basic conception of rater reliability including true score measurements of examinees’ language proficiency. This study was conducted based on a definition of rater reliability achieved by the reinterpretation of the various meanings of reliability. For these purposes, a basic framework of standardization was achieved using training theories, and this study proposes that a rater training program can be standardized by accomplishing innovative systematic changes that consider (a) the relevant literature, (b) the test instrument itself, (c) the test procedure, and (d) contextual effects such as the characteristics of the stakeholders, their concerns, the structure of the test, or washback effects of the test use. This study utilized a modified version of Lynch’s program evaluation model (1996; 2003) to collect evidence from different sources, including data drawn from the entire evaluation process ranging from needs analysis to a feedback system based on the final product of the evaluation. The effectiveness of both the training program and the individual performances were identified by incorporating all sources of data collected using measurement theory. Mixed methods were proposed for the data analysis. The data analysis involved an investigation of training effectiveness by measuring raters’ scoring reliability, and providing a new training program for raters’ professional improvement. Quantitative data analysis was proposed for analyzing the surveys, the rating corpus, and training effectiveness. Qualitative and document analysis were also essential for analyzing relevant training materials and workshop observation as well as exploring the degree of change in the perceptions of the raters. The results of this study provide educational implications for language testing. At the program level, standardized training contributed to shared responsibilities among test users. The results support the idea that the professionalism of the raters could be improved by providing access to similar quality input which can reinforce their learning and skills via training. The salient value of this dissertation is the collaboration with stakeholders in a test administration situation. Stakeholders’ concerns and challenges were clearly identified, shared, and resolved with the practitioners (the EPT trainer and raters). In addition, I recognize the importance of a balance between understanding fundamental theoretical underpinnings and applying theory through practical experience. It could be concluded that this study contributes to the enhancement of rating validity and the cumulative growth in scoring reliability, as well as a positive washback effect for the future rater training program.","abstract_has_math":false,"creators":["Jang, So Young"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Educational Psychology","degree_department":null,"school":null,"contributors":["Davidson, Frederick G.","Chang, Hua-Hua","Zhang, Jinming","Sadler, Randall W."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2010,"date_issued":"2010-05-14T20:52:01Z","date_published":"2010-05-14T20:52:01Z","updated_at":"2026-07-22T22:25:08Z","subjects":["Rater training","Development of rater training program"],"languages":["en"],"rights":["Copyright 2010 So Young Jang"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/15599","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Davidson, Frederick G.","Chang, Hua-Hua","Zhang, Jinming","Sadler, Randall W."]},{"key":"dc:creator","label":"Author","values":["Jang, So Young"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2010-05-14T20:52:01Z","2012-05-15T10:00:50Z","2010-5"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Educational Psychology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Rater training","Development of rater training program"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2010 So Young Jang"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/15599"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The primary purposes of this study are to identify the characteristics of modeling a rater training program and to develop an efficient training model at the University of Illinois at Urbana-Champaign. This study focuses substantially on a basic conception of rater reliability including true score measurements of examinees’ language proficiency. This study was conducted based on a definition of rater reliability achieved by the reinterpretation of the various meanings of reliability. For these purposes, a basic framework of standardization was achieved using training theories, and this study proposes that a rater training program can be standardized by accomplishing innovative systematic changes that consider (a) the relevant literature, (b) the test instrument itself, (c) the test procedure, and (d) contextual effects such as the characteristics of the stakeholders, their concerns, the structure of the test, or washback effects of the test use. This study utilized a modified version of Lynch’s program evaluation model (1996; 2003) to collect evidence from different sources, including data drawn from the entire evaluation process ranging from needs analysis to a feedback system based on the final product of the evaluation. The effectiveness of both the training program and the individual performances were identified by incorporating all sources of data collected using measurement theory. Mixed methods were proposed for the data analysis. The data analysis involved an investigation of training effectiveness by measuring raters’ scoring reliability, and providing a new training program for raters’ professional improvement. Quantitative data analysis was proposed for analyzing the surveys, the rating corpus, and training effectiveness. Qualitative and document analysis were also essential for analyzing relevant training materials and workshop observation as well as exploring the degree of change in the perceptions of the raters. The results of this study provide educational implications for language testing. At the program level, standardized training contributed to shared responsibilities among test users. The results support the idea that the professionalism of the raters could be improved by providing access to similar quality input which can reinforce their learning and skills via training. The salient value of this dissertation is the collaboration with stakeholders in a test administration situation. Stakeholders’ concerns and challenges were clearly identified, shared, and resolved with the practitioners (the EPT trainer and raters). In addition, I recognize the importance of a balance between understanding fundamental theoretical underpinnings and applying theory through practical experience. It could be concluded that this study contributes to the enhancement of rating validity and the cumulative growth in scoring reliability, as well as a positive washback effect for the future rater training program.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-04-22T13:04:23Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Jang_So Young.doc: 2795008 bytes, checksum: 8957a690e6f8740821655b189eb02b6e (MD5) Jang_ So Young.pdf: 1444039 bytes, checksum: 21cb954565fdccbf1ed9d5cf5eeee26e (MD5)","Made available in DSpace on 2010-05-14T20:52:01Z (GMT). No. of bitstreams: 5 Jang_So Young.pdf: 1438463 bytes, checksum: 46caf9e067060320fc9af5812924261c (MD5) Jang_ So Young.pdf: 1444039 bytes, checksum: 21cb954565fdccbf1ed9d5cf5eeee26e (MD5) Jang_So Young.doc: 2791936 bytes, checksum: 7cbd2f7d7e8857f56a33d1ecfc64f4a0 (MD5) 1_Jang_So Young.pdf: 1437041 bytes, checksum: e235e067fa9d118facb3d9e46ca07955 (MD5) license.txt: 4055 bytes, checksum: bf08cbb222f8f2281b01f7eb61ddaae9 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2010-05-14T20:52:51Z Item is restricted until 2012-05-14T20:52:43Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2012-05-15T10:00:50Z Item was in collections: University of Illinois Dissertations and Theses (ID: 204) College of Education Dissertations and Theses (ID: 434) No. of bitstreams: 6 1_Jang_So Young.pdf.txt: 511819 bytes, checksum: 475587cc310db59ad8443c6e10232d13 (MD5) Jang_So Young.pdf: 1438463 bytes, checksum: 46caf9e067060320fc9af5812924261c (MD5) Jang_ So Young.pdf: 1444039 bytes, checksum: 21cb954565fdccbf1ed9d5cf5eeee26e (MD5) Jang_So Young.doc: 2791936 bytes, checksum: 7cbd2f7d7e8857f56a33d1ecfc64f4a0 (MD5) 1_Jang_So Young.pdf: 1437041 bytes, checksum: e235e067fa9d118facb3d9e46ca07955 (MD5) license.txt: 4055 bytes, checksum: bf08cbb222f8f2281b01f7eb61ddaae9 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2012-05-15T10:00:50Z"]},{"key":"dc:title","label":"Title","values":["The development and evaluation of a systematic training program for increasing both rater reliability and rating accuracy"]}]}],"canonical_facts":{"dc:contributor":["Davidson, Frederick G.","Chang, Hua-Hua","Zhang, Jinming","Sadler, Randall W."],"dc:creator":["Jang, So Young"],"dc:date":["2010-05-14T20:52:01Z","2012-05-15T10:00:50Z","2010-5"],"dc:description":["The primary purposes of this study are to identify the characteristics of modeling a rater training program and to develop an efficient training model at the University of Illinois at Urbana-Champaign. This study focuses substantially on a basic conception of rater reliability including true score measurements of examinees’ language proficiency. This study was conducted based on a definition of rater reliability achieved by the reinterpretation of the various meanings of reliability. For these purposes, a basic framework of standardization was achieved using training theories, and this study proposes that a rater training program can be standardized by accomplishing innovative systematic changes that consider (a) the relevant literature, (b) the test instrument itself, (c) the test procedure, and (d) contextual effects such as the characteristics of the stakeholders, their concerns, the structure of the test, or washback effects of the test use. This study utilized a modified version of Lynch’s program evaluation model (1996; 2003) to collect evidence from different sources, including data drawn from the entire evaluation process ranging from needs analysis to a feedback system based on the final product of the evaluation. The effectiveness of both the training program and the individual performances were identified by incorporating all sources of data collected using measurement theory. Mixed methods were proposed for the data analysis. The data analysis involved an investigation of training effectiveness by measuring raters’ scoring reliability, and providing a new training program for raters’ professional improvement. Quantitative data analysis was proposed for analyzing the surveys, the rating corpus, and training effectiveness. Qualitative and document analysis were also essential for analyzing relevant training materials and workshop observation as well as exploring the degree of change in the perceptions of the raters. The results of this study provide educational implications for language testing. At the program level, standardized training contributed to shared responsibilities among test users. The results support the idea that the professionalism of the raters could be improved by providing access to similar quality input which can reinforce their learning and skills via training. The salient value of this dissertation is the collaboration with stakeholders in a test administration situation. Stakeholders’ concerns and challenges were clearly identified, shared, and resolved with the practitioners (the EPT trainer and raters). In addition, I recognize the importance of a balance between understanding fundamental theoretical underpinnings and applying theory through practical experience. It could be concluded that this study contributes to the enhancement of rating validity and the cumulative growth in scoring reliability, as well as a positive washback effect for the future rater training program.","Item withdrawn by Mark Zulauf (zulauf@illinois.edu) on 2010-04-22T13:04:23Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Jang_So Young.doc: 2795008 bytes, checksum: 8957a690e6f8740821655b189eb02b6e (MD5) Jang_ So Young.pdf: 1444039 bytes, checksum: 21cb954565fdccbf1ed9d5cf5eeee26e (MD5)","Made available in DSpace on 2010-05-14T20:52:01Z (GMT). No. of bitstreams: 5 Jang_So Young.pdf: 1438463 bytes, checksum: 46caf9e067060320fc9af5812924261c (MD5) Jang_ So Young.pdf: 1444039 bytes, checksum: 21cb954565fdccbf1ed9d5cf5eeee26e (MD5) Jang_So Young.doc: 2791936 bytes, checksum: 7cbd2f7d7e8857f56a33d1ecfc64f4a0 (MD5) 1_Jang_So Young.pdf: 1437041 bytes, checksum: e235e067fa9d118facb3d9e46ca07955 (MD5) license.txt: 4055 bytes, checksum: bf08cbb222f8f2281b01f7eb61ddaae9 (MD5)","Item marked as restricted to the 'UIUC Users [automated]' Group (id=2) by William Ingram (wingram2@illinois.edu) on 2010-05-14T20:52:51Z Item is restricted until 2012-05-14T20:52:43Z","Item reinstated by Sarah Shreeves (sshreeve@illinois.edu) on 2012-05-15T10:00:50Z Item was in collections: University of Illinois Dissertations and Theses (ID: 204) College of Education Dissertations and Theses (ID: 434) No. of bitstreams: 6 1_Jang_So Young.pdf.txt: 511819 bytes, checksum: 475587cc310db59ad8443c6e10232d13 (MD5) Jang_So Young.pdf: 1438463 bytes, checksum: 46caf9e067060320fc9af5812924261c (MD5) Jang_ So Young.pdf: 1444039 bytes, checksum: 21cb954565fdccbf1ed9d5cf5eeee26e (MD5) Jang_So Young.doc: 2791936 bytes, checksum: 7cbd2f7d7e8857f56a33d1ecfc64f4a0 (MD5) 1_Jang_So Young.pdf: 1437041 bytes, checksum: e235e067fa9d118facb3d9e46ca07955 (MD5) license.txt: 4055 bytes, checksum: bf08cbb222f8f2281b01f7eb61ddaae9 (MD5)","Item released from any restrictions by Sarah Shreeves (sshreeve@illinois.edu) on 2012-05-15T10:00:50Z"],"dc:identifier":["http://hdl.handle.net/2142/15599"],"dc:language":["en"],"dc:rights":["Copyright 2010 So Young Jang"],"dc:subject":["Rater training","Development of rater training program"],"dc:title":["The development and evaluation of a systematic training program for increasing both rater reliability and rating accuracy"],"thesis:degree_discipline":["Educational Psychology"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:08Z"}