{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/101148"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/101148","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Comparison of four stopping rules in computerized adaptive testing and examination of their application to on-the-fly multistage testing","abstract":"Computerized adaptive testing (CAT) is a powerful and efficient approach in educational testing for both estimating ability and classifying examinees into groups. When the purpose is to classify students as either proficient or not proficient in ability, an accurate estimate is not necessary, and the test can stop whenever a satisfactory decision can be made. Therefore, the stopping rule is a critical element in variable-length adaptive testing. In this study, the efficiency of four stopping rules was compared in variable-length CAT designs (vl-CAT): ability confidence interval (ACI), sequential probability ratio test (SPRT), generalized likelihood ratio (GLR), and the truncation rule. In addition, their application to the newly-developed adaptive testing design, on-the-fly multistage testing (OMST), was also examined and compared with vl-CAT. Two simulation studies were conducted. In study 1, since the fourth stopping rule cannot be executed independently, ACI, SPRT, and GLR were combined with the truncation rule, which resulted in 6 CAT and 6 OMST designs in total. With the classification accuracy (CA) controlled at the same level, the test length of 12 variable-length designs was examined. All test designs in study 1 have a length between 10 and 30. In study 2, the lower and upper bound of the test length was extended to 30 and 100, and only ACI- and GLR- CAT designs were conducted to provide a more general comparison of these two stopping rules. In both studies, the ability was estimated by maximum likelihood estimation or expected a posterior. The next item(s) is (are) selected with the maximum priority index at the current ability estimate. 1000 theta values were simulated from a standard normal distribution, and 30 replications were conducted under each design. The results show that OMST produced similar results to CAT. Regarding the efficiency of four stopping rules, the truncated versions of ACI, SPRT, and GLR produced shorter test lengths than their corresponding counterparts. Among ACI, SPRT, and GLR, SPRT yielded the longest test length with the highest estimation accuracy. The results of GLR and ACI designs are similar, but ACI is more efficient for examinees whose ability is far from the cutoff point, and GLR is more efficient for examinees whose ability is near the cutoff point. It can be concluded that the stopping rules designed for CAT also function for OMST in a similar way. When the item selection method is estimate-based rather than cutscore-based, SPRT performed less efficiently than ACI and GLR. The efficiency of GLR and ACI is comparable, and each has its own strengths. The truncation rule is useful because it prevents examinees from taking unnecessary items. These results imply the good statistical properties of variable-length OMST, which facilitates its future application. The research also provides a direct comparison between different stopping rules, giving practitioners more information about the above-mentioned applicable situations of different rules. The studies also develop a new simple truncation rule and indicate its feasibility in a future adaptive testing context.","abstract_html":"Computerized adaptive testing (CAT) is a powerful and efficient approach in educational testing for both estimating ability and classifying examinees into groups. When the purpose is to classify students as either proficient or not proficient in ability, an accurate estimate is not necessary, and the test can stop whenever a satisfactory decision can be made. Therefore, the stopping rule is a critical element in variable-length adaptive testing. In this study, the efficiency of four stopping rules was compared in variable-length CAT designs (vl-CAT): ability confidence interval (ACI), sequential probability ratio test (SPRT), generalized likelihood ratio (GLR), and the truncation rule. In addition, their application to the newly-developed adaptive testing design, on-the-fly multistage testing (OMST), was also examined and compared with vl-CAT. Two simulation studies were conducted. In study 1, since the fourth stopping rule cannot be executed independently, ACI, SPRT, and GLR were combined with the truncation rule, which resulted in 6 CAT and 6 OMST designs in total. With the classification accuracy (CA) controlled at the same level, the test length of 12 variable-length designs was examined. All test designs in study 1 have a length between 10 and 30. In study 2, the lower and upper bound of the test length was extended to 30 and 100, and only ACI- and GLR- CAT designs were conducted to provide a more general comparison of these two stopping rules. In both studies, the ability was estimated by maximum likelihood estimation or expected a posterior. The next item(s) is (are) selected with the maximum priority index at the current ability estimate. 1000 theta values were simulated from a standard normal distribution, and 30 replications were conducted under each design. The results show that OMST produced similar results to CAT. Regarding the efficiency of four stopping rules, the truncated versions of ACI, SPRT, and GLR produced shorter test lengths than their corresponding counterparts. Among ACI, SPRT, and GLR, SPRT yielded the longest test length with the highest estimation accuracy. The results of GLR and ACI designs are similar, but ACI is more efficient for examinees whose ability is far from the cutoff point, and GLR is more efficient for examinees whose ability is near the cutoff point. It can be concluded that the stopping rules designed for CAT also function for OMST in a similar way. When the item selection method is estimate-based rather than cutscore-based, SPRT performed less efficiently than ACI and GLR. The efficiency of GLR and ACI is comparable, and each has its own strengths. The truncation rule is useful because it prevents examinees from taking unnecessary items. These results imply the good statistical properties of variable-length OMST, which facilitates its future application. The research also provides a direct comparison between different stopping rules, giving practitioners more information about the above-mentioned applicable situations of different rules. The studies also develop a new simple truncation rule and indicate its feasibility in a future adaptive testing context.","abstract_has_math":false,"creators":["Tian, Chen"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Educational Psychology","degree_department":null,"school":null,"contributors":["Chang, Hua-Hua","Anderson, Carolyn","Zhang, Jinming"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2018,"date_issued":"2018-09-04T20:34:01Z","date_published":"2018-09-04T20:34:01Z","updated_at":"2026-07-22T22:24:38Z","subjects":["variable-length on-the-fly multistage testing","stopping rules","computerized adaptive testing","mastery testing"],"languages":["en"],"rights":["Copyright 2018 Chen Tian"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/101148","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Chang, Hua-Hua","Anderson, Carolyn","Zhang, Jinming"]},{"key":"dc:creator","label":"Author","values":["Tian, Chen"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2018-09-04T20:34:01Z","2020-09-05T09:15:16Z","2018-04-16","2018-05"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Educational Psychology"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["variable-length on-the-fly multistage testing","stopping rules","computerized adaptive testing","mastery testing"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2018 Chen Tian"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/101148"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Computerized adaptive testing (CAT) is a powerful and efficient approach in educational testing for both estimating ability and classifying examinees into groups. When the purpose is to classify students as either proficient or not proficient in ability, an accurate estimate is not necessary, and the test can stop whenever a satisfactory decision can be made. Therefore, the stopping rule is a critical element in variable-length adaptive testing. In this study, the efficiency of four stopping rules was compared in variable-length CAT designs (vl-CAT): ability confidence interval (ACI), sequential probability ratio test (SPRT), generalized likelihood ratio (GLR), and the truncation rule. In addition, their application to the newly-developed adaptive testing design, on-the-fly multistage testing (OMST), was also examined and compared with vl-CAT. Two simulation studies were conducted. In study 1, since the fourth stopping rule cannot be executed independently, ACI, SPRT, and GLR were combined with the truncation rule, which resulted in 6 CAT and 6 OMST designs in total. With the classification accuracy (CA) controlled at the same level, the test length of 12 variable-length designs was examined. All test designs in study 1 have a length between 10 and 30. In study 2, the lower and upper bound of the test length was extended to 30 and 100, and only ACI- and GLR- CAT designs were conducted to provide a more general comparison of these two stopping rules. In both studies, the ability was estimated by maximum likelihood estimation or expected a posterior. The next item(s) is (are) selected with the maximum priority index at the current ability estimate. 1000 theta values were simulated from a standard normal distribution, and 30 replications were conducted under each design. The results show that OMST produced similar results to CAT. Regarding the efficiency of four stopping rules, the truncated versions of ACI, SPRT, and GLR produced shorter test lengths than their corresponding counterparts. Among ACI, SPRT, and GLR, SPRT yielded the longest test length with the highest estimation accuracy. The results of GLR and ACI designs are similar, but ACI is more efficient for examinees whose ability is far from the cutoff point, and GLR is more efficient for examinees whose ability is near the cutoff point. It can be concluded that the stopping rules designed for CAT also function for OMST in a similar way. When the item selection method is estimate-based rather than cutscore-based, SPRT performed less efficiently than ACI and GLR. The efficiency of GLR and ACI is comparable, and each has its own strengths. The truncation rule is useful because it prevents examinees from taking unnecessary items. These results imply the good statistical properties of variable-length OMST, which facilitates its future application. The research also provides a direct comparison between different stopping rules, giving practitioners more information about the above-mentioned applicable situations of different rules. The studies also develop a new simple truncation rule and indicate its feasibility in a future adaptive testing context.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-05-01","The student, Chen Tian, accepted the attached license on 2018-04-10 at 20:52.","The student, Chen Tian, submitted this Thesis for approval on 2018-04-11 at 19:59.","This Thesis was approved for publication on 2018-04-16 at 08:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12183 on 2018-08-31 at 17:18:23","Made available in DSpace on 2018-09-04T20:34:01Z (GMT). No. of bitstreams: 2 TIAN-THESIS-2018.pdf: 1812644 bytes, checksum: bb0f2400d72959bb3878eafac472c9ab (MD5) LICENSE.txt: 4206 bytes, checksum: 08b8da6442603a7ec8a16b40ed6a30fe (MD5) Previous issue date: 2018-04-16","Embargo set by: Seth Robbins for item 107231 Lift date: 2020-09-04T20:34:13Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107231 Lift date: 2020-09-04T20:37:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107231 Lift date: 2020-09-04T20:42:08Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107231 on 2020-09-05T09:15:16Z."]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Comparison of four stopping rules in computerized adaptive testing and examination of their application to on-the-fly multistage testing"]}]}],"canonical_facts":{"dc:contributor":["Chang, Hua-Hua","Anderson, Carolyn","Zhang, Jinming"],"dc:creator":["Tian, Chen"],"dc:date":["2018-09-04T20:34:01Z","2020-09-05T09:15:16Z","2018-04-16","2018-05"],"dc:description":["Computerized adaptive testing (CAT) is a powerful and efficient approach in educational testing for both estimating ability and classifying examinees into groups. When the purpose is to classify students as either proficient or not proficient in ability, an accurate estimate is not necessary, and the test can stop whenever a satisfactory decision can be made. Therefore, the stopping rule is a critical element in variable-length adaptive testing. In this study, the efficiency of four stopping rules was compared in variable-length CAT designs (vl-CAT): ability confidence interval (ACI), sequential probability ratio test (SPRT), generalized likelihood ratio (GLR), and the truncation rule. In addition, their application to the newly-developed adaptive testing design, on-the-fly multistage testing (OMST), was also examined and compared with vl-CAT. Two simulation studies were conducted. In study 1, since the fourth stopping rule cannot be executed independently, ACI, SPRT, and GLR were combined with the truncation rule, which resulted in 6 CAT and 6 OMST designs in total. With the classification accuracy (CA) controlled at the same level, the test length of 12 variable-length designs was examined. All test designs in study 1 have a length between 10 and 30. In study 2, the lower and upper bound of the test length was extended to 30 and 100, and only ACI- and GLR- CAT designs were conducted to provide a more general comparison of these two stopping rules. In both studies, the ability was estimated by maximum likelihood estimation or expected a posterior. The next item(s) is (are) selected with the maximum priority index at the current ability estimate. 1000 theta values were simulated from a standard normal distribution, and 30 replications were conducted under each design. The results show that OMST produced similar results to CAT. Regarding the efficiency of four stopping rules, the truncated versions of ACI, SPRT, and GLR produced shorter test lengths than their corresponding counterparts. Among ACI, SPRT, and GLR, SPRT yielded the longest test length with the highest estimation accuracy. The results of GLR and ACI designs are similar, but ACI is more efficient for examinees whose ability is far from the cutoff point, and GLR is more efficient for examinees whose ability is near the cutoff point. It can be concluded that the stopping rules designed for CAT also function for OMST in a similar way. When the item selection method is estimate-based rather than cutscore-based, SPRT performed less efficiently than ACI and GLR. The efficiency of GLR and ACI is comparable, and each has its own strengths. The truncation rule is useful because it prevents examinees from taking unnecessary items. These results imply the good statistical properties of variable-length OMST, which facilitates its future application. The research also provides a direct comparison between different stopping rules, giving practitioners more information about the above-mentioned applicable situations of different rules. The studies also develop a new simple truncation rule and indicate its feasibility in a future adaptive testing context.","Submission published under a 24 month embargo labeled 'U of I Access', the embargo will last until 2020-05-01","The student, Chen Tian, accepted the attached license on 2018-04-10 at 20:52.","The student, Chen Tian, submitted this Thesis for approval on 2018-04-11 at 19:59.","This Thesis was approved for publication on 2018-04-16 at 08:49.","DSpace SAF Submission Ingestion Package generated from Vireo submission #12183 on 2018-08-31 at 17:18:23","Made available in DSpace on 2018-09-04T20:34:01Z (GMT). No. of bitstreams: 2 TIAN-THESIS-2018.pdf: 1812644 bytes, checksum: bb0f2400d72959bb3878eafac472c9ab (MD5) LICENSE.txt: 4206 bytes, checksum: 08b8da6442603a7ec8a16b40ed6a30fe (MD5) Previous issue date: 2018-04-16","Embargo set by: Seth Robbins for item 107231 Lift date: 2020-09-04T20:34:13Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107231 Lift date: 2020-09-04T20:37:00Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","Embargo set by: Seth Robbins for item 107231 Lift date: 2020-09-04T20:42:08Z Reason: Author requested U of Illinois access only (OA after 2yrs) in Vireo ETD system","U of I Only Restriction Lifted for Item 107231 on 2020-09-05T09:15:16Z."],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/101148"],"dc:language":["en"],"dc:rights":["Copyright 2018 Chen Tian"],"dc:subject":["variable-length on-the-fly multistage testing","stopping rules","computerized adaptive testing","mastery testing"],"dc:title":["Comparison of four stopping rules in computerized adaptive testing and examination of their application to on-the-fly multistage testing"],"dc:type":["text"],"thesis:degree_discipline":["Educational Psychology"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:38Z"}