{"id":{"repo_id":"manitoba","oai_identifier":"oai:mspace.lib.umanitoba.ca:1993/39496"},"canonical_url":"https://search.dev.ndltd.org/etd/manitoba/oai:mspace.lib.umanitoba.ca:1993/39496","repository":{"repo_id":"manitoba","name":"University of Manitoba","base_url":"https://mspace.lib.umanitoba.ca/oai/request"},"display":{"title":"Comparison of machine learning methods with an underlying logistic regression model to detect differential item functioning in patient-reported outcome measures","abstract":"Background: Patient-reported outcome measures (PROMs) are self-report instruments about health-related quality of life and well-being (i.e., physical and mental health). Assessing the validity of a PROM involves ensuring measurement invariance (MI), which confirms unbiased comparisons across groups. Differential item functioning (DIF), a form of measurement non-invariance, occurs when individuals with the same level of health status respond to an item differently. Many studies about DIF in PROMs focus on demographic characteristics (e.g., age), but other characteristics of individuals, such as the presence of comorbid health conditions, may also contribute to DIF. Machine learning (ML) methods may be advantageous to test for DIF across multiple covariates. The research purpose was to test for DIF using ML tree-based and penalized methods based on logistic regression (LR). The objectives were to 1) compare the performance of two ML methods to detect DIF on multiple covariates, and 2) test the association of demographic and clinical covariates with DIF. Methods: DIF was tested using an item-focused tree (IFT) with an underlying LR model and a Least Absolute Shrinkage and Selection Operator (LASSO) regression model. For Objective 1, a simulation study was conducted in which data were generated under different analytical conditions by varying sample sizes, DIF effect magnitudes, and correlations among covariates. The performance of both the IFT and LASSO regression models was evaluated using Type I error and statistical power rates. For Objective 2, the association of the covariates with DIF was assessed in the 36-item Short Form Survey items completed by individuals diagnosed with immune-mediated inflammatory diseases. A hold-out bootstrap cross-validation technique was conducted to evaluate the performance of the IFT and LASSO regression models using the Brier score, accuracy, and mean square error (MSE) in test data. Results: In the simulation study, the Type I error rate for the IFT regression model was below the nominal 5% level across simulation conditions. The LASSO regression model had lower Type I error rates than the IFT model when the correlation between covariates was strong. In conditions with small and moderate DIF effect sizes, the LASSO regression model had greater statistical power than the IFT regression model. In real-world data, the IFT regression model detected eight DIF items, and the LASSO regression model detected 16 DIF items. Sex and hypertension were the most frequent covariates associated with DIF. When both methods flagged an item for DIF, they identified similar associated covariates. The estimated Brier score, accuracy, and MSE were similar for both methods. Conclusions: Establishing MI in a PROM means that the latent construct is equivalent across population groups defined by socio-demographic and personal health characteristics. MI contributes to equity-focused PROM development. The IFT regression model is recommended to control the rate of false positives, where DIF is incorrectly detected. The LASSO regression model is recommended when DIF effects are small.","abstract_html":"Background: Patient-reported outcome measures (PROMs) are self-report instruments about health-related quality of life and well-being (i.e., physical and mental health). Assessing the validity of a PROM involves ensuring measurement invariance (MI), which confirms unbiased comparisons across groups. Differential item functioning (DIF), a form of measurement non-invariance, occurs when individuals with the same level of health status respond to an item differently. Many studies about DIF in PROMs focus on demographic characteristics (e.g., age), but other characteristics of individuals, such as the presence of comorbid health conditions, may also contribute to DIF. Machine learning (ML) methods may be advantageous to test for DIF across multiple covariates. The research purpose was to test for DIF using ML tree-based and penalized methods based on logistic regression (LR). The objectives were to 1) compare the performance of two ML methods to detect DIF on multiple covariates, and 2) test the association of demographic and clinical covariates with DIF. Methods: DIF was tested using an item-focused tree (IFT) with an underlying LR model and a Least Absolute Shrinkage and Selection Operator (LASSO) regression model. For Objective 1, a simulation study was conducted in which data were generated under different analytical conditions by varying sample sizes, DIF effect magnitudes, and correlations among covariates. The performance of both the IFT and LASSO regression models was evaluated using Type I error and statistical power rates. For Objective 2, the association of the covariates with DIF was assessed in the 36-item Short Form Survey items completed by individuals diagnosed with immune-mediated inflammatory diseases. A hold-out bootstrap cross-validation technique was conducted to evaluate the performance of the IFT and LASSO regression models using the Brier score, accuracy, and mean square error (MSE) in test data. Results: In the simulation study, the Type I error rate for the IFT regression model was below the nominal 5% level across simulation conditions. The LASSO regression model had lower Type I error rates than the IFT model when the correlation between covariates was strong. In conditions with small and moderate DIF effect sizes, the LASSO regression model had greater statistical power than the IFT regression model. In real-world data, the IFT regression model detected eight DIF items, and the LASSO regression model detected 16 DIF items. Sex and hypertension were the most frequent covariates associated with DIF. When both methods flagged an item for DIF, they identified similar associated covariates. The estimated Brier score, accuracy, and MSE were similar for both methods. Conclusions: Establishing MI in a PROM means that the latent construct is equivalent across population groups defined by socio-demographic and personal health characteristics. MI contributes to equity-focused PROM development. The IFT regression model is recommended to control the rate of false positives, where DIF is incorrectly detected. The LASSO regression model is recommended when DIF effects are small.","abstract_has_math":false,"creators":["Dorostimotlagh, Razieh"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Lix, Lisa"],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12-10","date_published":"2025-12-10","updated_at":"2026-08-21T22:21:56Z","subjects":["Differential item functioning","Quality of life","Measurement invariance","Item-focused tree","LASSO","Comorbid conditions","Post-covariate selection"],"languages":["eng"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1993/39496","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"source_record":{"url":"https://mspace.lib.umanitoba.ca/oai/request?verb=GetRecord&metadataPrefix=dim&identifier=oai%3Amspace.lib.umanitoba.ca%3A1993%2F39496","prefix":"dim"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.supervisor","label":"Supervisor","values":["Lix, Lisa"]},{"key":"dc:creator","label":"Author","values":["Dorostimotlagh, Razieh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2025-12-16T15:30:17Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2025-12-16T15:30:17Z"]},{"key":"dc:date.issued","label":"Date","values":["2025-12-10"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Differential item functioning","Quality of life","Measurement invariance","Item-focused tree","LASSO","Comorbid conditions","Post-covariate selection"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1993/39496"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Background: Patient-reported outcome measures (PROMs) are self-report instruments about health-related quality of life and well-being (i.e., physical and mental health). Assessing the validity of a PROM involves ensuring measurement invariance (MI), which confirms unbiased comparisons across groups. Differential item functioning (DIF), a form of measurement non-invariance, occurs when individuals with the same level of health status respond to an item differently. Many studies about DIF in PROMs focus on demographic characteristics (e.g., age), but other characteristics of individuals, such as the presence of comorbid health conditions, may also contribute to DIF. Machine learning (ML) methods may be advantageous to test for DIF across multiple covariates. The research purpose was to test for DIF using ML tree-based and penalized methods based on logistic regression (LR). The objectives were to 1) compare the performance of two ML methods to detect DIF on multiple covariates, and 2) test the association of demographic and clinical covariates with DIF. Methods: DIF was tested using an item-focused tree (IFT) with an underlying LR model and a Least Absolute Shrinkage and Selection Operator (LASSO) regression model. For Objective 1, a simulation study was conducted in which data were generated under different analytical conditions by varying sample sizes, DIF effect magnitudes, and correlations among covariates. The performance of both the IFT and LASSO regression models was evaluated using Type I error and statistical power rates. For Objective 2, the association of the covariates with DIF was assessed in the 36-item Short Form Survey items completed by individuals diagnosed with immune-mediated inflammatory diseases. A hold-out bootstrap cross-validation technique was conducted to evaluate the performance of the IFT and LASSO regression models using the Brier score, accuracy, and mean square error (MSE) in test data. Results: In the simulation study, the Type I error rate for the IFT regression model was below the nominal 5% level across simulation conditions. The LASSO regression model had lower Type I error rates than the IFT model when the correlation between covariates was strong. In conditions with small and moderate DIF effect sizes, the LASSO regression model had greater statistical power than the IFT regression model. In real-world data, the IFT regression model detected eight DIF items, and the LASSO regression model detected 16 DIF items. Sex and hypertension were the most frequent covariates associated with DIF. When both methods flagged an item for DIF, they identified similar associated covariates. The estimated Brier score, accuracy, and MSE were similar for both methods. Conclusions: Establishing MI in a PROM means that the latent construct is equivalent across population groups defined by socio-demographic and personal health characteristics. MI contributes to equity-focused PROM development. The IFT regression model is recommended to control the rate of false positives, where DIF is incorrectly detected. The LASSO regression model is recommended when DIF effects are small."]},{"key":"dc:title","label":"Title","values":["Comparison of machine learning methods with an underlying logistic regression model to detect differential item functioning in patient-reported outcome measures"]}]}],"canonical_facts":{"dc:contributor.supervisor":["Lix, Lisa"],"dc:creator":["Dorostimotlagh, Razieh"],"dc:date.accessioned":["2025-12-16T15:30:17Z"],"dc:date.available":["2025-12-16T15:30:17Z"],"dc:date.issued":["2025-12-10"],"dc:description.abstract":["Background: Patient-reported outcome measures (PROMs) are self-report instruments about health-related quality of life and well-being (i.e., physical and mental health). Assessing the validity of a PROM involves ensuring measurement invariance (MI), which confirms unbiased comparisons across groups. Differential item functioning (DIF), a form of measurement non-invariance, occurs when individuals with the same level of health status respond to an item differently. Many studies about DIF in PROMs focus on demographic characteristics (e.g., age), but other characteristics of individuals, such as the presence of comorbid health conditions, may also contribute to DIF. Machine learning (ML) methods may be advantageous to test for DIF across multiple covariates. The research purpose was to test for DIF using ML tree-based and penalized methods based on logistic regression (LR). The objectives were to 1) compare the performance of two ML methods to detect DIF on multiple covariates, and 2) test the association of demographic and clinical covariates with DIF. Methods: DIF was tested using an item-focused tree (IFT) with an underlying LR model and a Least Absolute Shrinkage and Selection Operator (LASSO) regression model. For Objective 1, a simulation study was conducted in which data were generated under different analytical conditions by varying sample sizes, DIF effect magnitudes, and correlations among covariates. The performance of both the IFT and LASSO regression models was evaluated using Type I error and statistical power rates. For Objective 2, the association of the covariates with DIF was assessed in the 36-item Short Form Survey items completed by individuals diagnosed with immune-mediated inflammatory diseases. A hold-out bootstrap cross-validation technique was conducted to evaluate the performance of the IFT and LASSO regression models using the Brier score, accuracy, and mean square error (MSE) in test data. Results: In the simulation study, the Type I error rate for the IFT regression model was below the nominal 5% level across simulation conditions. The LASSO regression model had lower Type I error rates than the IFT model when the correlation between covariates was strong. In conditions with small and moderate DIF effect sizes, the LASSO regression model had greater statistical power than the IFT regression model. In real-world data, the IFT regression model detected eight DIF items, and the LASSO regression model detected 16 DIF items. Sex and hypertension were the most frequent covariates associated with DIF. When both methods flagged an item for DIF, they identified similar associated covariates. The estimated Brier score, accuracy, and MSE were similar for both methods. Conclusions: Establishing MI in a PROM means that the latent construct is equivalent across population groups defined by socio-demographic and personal health characteristics. MI contributes to equity-focused PROM development. The IFT regression model is recommended to control the rate of false positives, where DIF is incorrectly detected. The LASSO regression model is recommended when DIF effects are small."],"dc:identifier.uri":["http://hdl.handle.net/1993/39496"],"dc:language.iso":["eng"],"dc:subject":["Differential item functioning","Quality of life","Measurement invariance","Item-focused tree","LASSO","Comorbid conditions","Post-covariate selection"],"dc:title":["Comparison of machine learning methods with an underlying logistic regression model to detect differential item functioning in patient-reported outcome measures"]},"updated_at":"2026-08-21T22:21:56Z"}