{"id":{"repo_id":"wayne-thes","oai_identifier":"oai:digitalcommons.wayne.edu:oa_dissertations-2237"},"canonical_url":"https://search.dev.ndltd.org/etd/wayne-thes/oai:digitalcommons.wayne.edu:oa_dissertations-2237","repository":{"repo_id":"wayne-thes","name":"Wayne State University","base_url":"https://digitalcommons.wayne.edu/do/oai/"},"display":{"title":"An evaluation of item response theory for detecting differential item functioning of examinees' responses to the seventh grade mathematics MEAP test investigating learner characteristics /","abstract":"Historically test fairness has been an issue for stakeholders such as test makers, administrators, educators, examinees, and others. The consequence of bias testing creates \"high stakes\" for all stakeholders. During the past 50 years several cycles of dissatisfaction and reform occurred in educational testing (McGhee, 1995). Large scale testing was used to make decisions regarding college entrance, employment opportunities, funding, school policies and curriculum. The Michigan Educational Assessment Program (MEAP) was a large scale testing program requiring students to pass several specific sections (reading, mathematics, and science) before receiving a state-endorsed diploma. Test measurement continues to come under the microscope because of the \"high stake\" outcomes. Early on classical statistics were used which experienced several short comings. Therefore, item response theory was developed in an attempt to explain difference in observed responses and unobserved characteristics such as ability. Statistical techniques were developed to detect differential item functioning (DIF) between two groups. This study measured DIF between gender and race which used the Mantel-Haenszel test statistic and the three parameters logistical model. Both techniques detected DIF but not for the same test items. The Mantel-Haenszel was an easier statistical method to use, with approximately 23 (20%) of the test items flagged for gender, and 19 (16.5%) flagged for race. The three parameter logistical model flagged 15 (13%) for gender, and 18 (15.6%) for race. This model was difficult to use in its DOS extended version however, it provided more information to the user. The difficulty with differential item functioning (DIF) was once the test item had been flagged the results would need to be carefully interpreted to determine if item bias existed. Although both techniques failed to detect the same test items, both tended to support the literature in terms of gender and race difference for mathematical test items. Further examination of the differences in statistical test for significance in an effort to replicate the findings.","abstract_html":"Historically test fairness has been an issue for stakeholders such as test makers, administrators, educators, examinees, and others. The consequence of bias testing creates &quot;high stakes&quot; for all stakeholders. During the past 50 years several cycles of dissatisfaction and reform occurred in educational testing (McGhee, 1995). Large scale testing was used to make decisions regarding college entrance, employment opportunities, funding, school policies and curriculum. The Michigan Educational Assessment Program (MEAP) was a large scale testing program requiring students to pass several specific sections (reading, mathematics, and science) before receiving a state-endorsed diploma. Test measurement continues to come under the microscope because of the &quot;high stake&quot; outcomes. Early on classical statistics were used which experienced several short comings. Therefore, item response theory was developed in an attempt to explain difference in observed responses and unobserved characteristics such as ability. Statistical techniques were developed to detect differential item functioning (DIF) between two groups. This study measured DIF between gender and race which used the Mantel-Haenszel test statistic and the three parameters logistical model. Both techniques detected DIF but not for the same test items. The Mantel-Haenszel was an easier statistical method to use, with approximately 23 (20%) of the test items flagged for gender, and 19 (16.5%) flagged for race. The three parameter logistical model flagged 15 (13%) for gender, and 18 (15.6%) for race. This model was difficult to use in its DOS extended version however, it provided more information to the user. The difficulty with differential item functioning (DIF) was once the test item had been flagged the results would need to be carefully interpreted to determine if item bias existed. Although both techniques failed to detect the same test items, both tended to support the literature in terms of gender and race difference for mathematical test items. Further examination of the differences in statistical test for significance in an effort to replicate the findings.","abstract_has_math":false,"creators":["Odett, David Charles"],"institution":null,"degree_name":"Ph.D.","degree_level":"Open Access Dissertation","degree_discipline":"Education Evaluation and Research","degree_department":null,"school":null,"contributors":["Donald R. Marcotte, Ph.D."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":1997,"date_issued":"1997-01-01T08:00:00Z","date_published":"1997-01-01T08:00:00Z","updated_at":"2026-07-24T06:00:09Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://digitalcommons.wayne.edu/oa_dissertations/1238","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Donald R. Marcotte, Ph.D."]},{"key":"dc:creator","label":"Author","values":["Odett, David Charles"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2015-10-01T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Education Evaluation and Research"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Open Access Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://digitalcommons.wayne.edu/oa_dissertations/1238"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Historically test fairness has been an issue for stakeholders such as test makers, administrators, educators, examinees, and others. The consequence of bias testing creates \"high stakes\" for all stakeholders. During the past 50 years several cycles of dissatisfaction and reform occurred in educational testing (McGhee, 1995). Large scale testing was used to make decisions regarding college entrance, employment opportunities, funding, school policies and curriculum. The Michigan Educational Assessment Program (MEAP) was a large scale testing program requiring students to pass several specific sections (reading, mathematics, and science) before receiving a state-endorsed diploma. Test measurement continues to come under the microscope because of the \"high stake\" outcomes. Early on classical statistics were used which experienced several short comings. Therefore, item response theory was developed in an attempt to explain difference in observed responses and unobserved characteristics such as ability. Statistical techniques were developed to detect differential item functioning (DIF) between two groups. This study measured DIF between gender and race which used the Mantel-Haenszel test statistic and the three parameters logistical model. Both techniques detected DIF but not for the same test items. The Mantel-Haenszel was an easier statistical method to use, with approximately 23 (20%) of the test items flagged for gender, and 19 (16.5%) flagged for race. The three parameter logistical model flagged 15 (13%) for gender, and 18 (15.6%) for race. This model was difficult to use in its DOS extended version however, it provided more information to the user. The difficulty with differential item functioning (DIF) was once the test item had been flagged the results would need to be carefully interpreted to determine if item bias existed. Although both techniques failed to detect the same test items, both tended to support the literature in terms of gender and race difference for mathematical test items. Further examination of the differences in statistical test for significance in an effort to replicate the findings."]},{"key":"dc:title","label":"Title","values":["An evaluation of item response theory for detecting differential item functioning of examinees' responses to the seventh grade mathematics MEAP test investigating learner characteristics /"]}]}],"canonical_facts":{"dc:contributor":["Donald R. Marcotte, Ph.D."],"dc:creator":["Odett, David Charles"],"dc:date.available":["2015-10-01T07:00:00Z"],"dc:description.abstract":["Historically test fairness has been an issue for stakeholders such as test makers, administrators, educators, examinees, and others. The consequence of bias testing creates \"high stakes\" for all stakeholders. During the past 50 years several cycles of dissatisfaction and reform occurred in educational testing (McGhee, 1995). Large scale testing was used to make decisions regarding college entrance, employment opportunities, funding, school policies and curriculum. The Michigan Educational Assessment Program (MEAP) was a large scale testing program requiring students to pass several specific sections (reading, mathematics, and science) before receiving a state-endorsed diploma. Test measurement continues to come under the microscope because of the \"high stake\" outcomes. Early on classical statistics were used which experienced several short comings. Therefore, item response theory was developed in an attempt to explain difference in observed responses and unobserved characteristics such as ability. Statistical techniques were developed to detect differential item functioning (DIF) between two groups. This study measured DIF between gender and race which used the Mantel-Haenszel test statistic and the three parameters logistical model. Both techniques detected DIF but not for the same test items. The Mantel-Haenszel was an easier statistical method to use, with approximately 23 (20%) of the test items flagged for gender, and 19 (16.5%) flagged for race. The three parameter logistical model flagged 15 (13%) for gender, and 18 (15.6%) for race. This model was difficult to use in its DOS extended version however, it provided more information to the user. The difficulty with differential item functioning (DIF) was once the test item had been flagged the results would need to be carefully interpreted to determine if item bias existed. Although both techniques failed to detect the same test items, both tended to support the literature in terms of gender and race difference for mathematical test items. Further examination of the differences in statistical test for significance in an effort to replicate the findings."],"dc:identifier":["https://digitalcommons.wayne.edu/oa_dissertations/1238"],"dc:title":["An evaluation of item response theory for detecting differential item functioning of examinees' responses to the seventh grade mathematics MEAP test investigating learner characteristics /"],"thesis:degree_discipline":["Education Evaluation and Research"],"thesis:degree_level":["Open Access Dissertation"],"thesis:degree_name":["Ph.D."]},"updated_at":"2026-07-24T06:00:09Z"}