{"id":{"repo_id":"york","oai_identifier":"oai:yorkspace.library.yorku.ca:10315/42143"},"canonical_url":"https://search.dev.ndltd.org/etd/york/oai:yorkspace.library.yorku.ca:10315/42143","repository":{"repo_id":"york","name":"York University","base_url":"https://yorkspace.library.yorku.ca/oai/request"},"display":{"title":"A First Look at Fairness of Machine Learning Based Code Reviewer Recommendation","abstract":"The fairness of machine learning (ML) approaches is critical to the reliability of modern artificial intelligence systems. Despite extensive study on this topic, the fairness of ML models in the software engineering (SE) domain has yet to be well explored. As a result, many ML-powered software systems, particularly those utilized in the software engineering community, continue to be prone to fairness issues. Taking one of the typical SE tasks, i.e., code reviewer recommendation, as a subject, this work conducts the first study toward investigating the fairness of ML applications in the SE domain, explicitly focusing on the code reviewer recommendation task. Our empirical study demonstrates that current state-of-the-art ML-based code reviewer recommendation techniques exhibit unfairness and discriminating behaviors. Specifically, male reviewers get, on average, 7.25% more recommendations than female code reviewers compared to their distribution in the reviewer set. This work also discusses why the studied ML-based code reviewer recommendation systems are unfair and provides solutions to mitigate the unfairness. For instance, these techniques may recommend male reviewers at a significantly higher rate than female reviewers in a discriminatory manner. Our study further indicates that existing mitigation methods can significantly enhance fairness in projects with a similar distribution of protected and privileged groups. Still, their effectiveness in improving fairness on imbalanced or skewed data is limited. Eventually, we suggest a solution to overcome the drawbacks of existing mitigation techniques and tackle bias in imbalanced or skewed datasets.","abstract_html":"The fairness of machine learning (ML) approaches is critical to the reliability of modern artificial intelligence systems. Despite extensive study on this topic, the fairness of ML models in the software engineering (SE) domain has yet to be well explored. As a result, many ML-powered software systems, particularly those utilized in the software engineering community, continue to be prone to fairness issues. Taking one of the typical SE tasks, i.e., code reviewer recommendation, as a subject, this work conducts the first study toward investigating the fairness of ML applications in the SE domain, explicitly focusing on the code reviewer recommendation task. Our empirical study demonstrates that current state-of-the-art ML-based code reviewer recommendation techniques exhibit unfairness and discriminating behaviors. Specifically, male reviewers get, on average, 7.25% more recommendations than female code reviewers compared to their distribution in the reviewer set. This work also discusses why the studied ML-based code reviewer recommendation systems are unfair and provides solutions to mitigate the unfairness. For instance, these techniques may recommend male reviewers at a significantly higher rate than female reviewers in a discriminatory manner. Our study further indicates that existing mitigation methods can significantly enhance fairness in projects with a similar distribution of protected and privileged groups. Still, their effectiveness in improving fairness on imbalanced or skewed data is limited. Eventually, we suggest a solution to overcome the drawbacks of existing mitigation techniques and tackle bias in imbalanced or skewed datasets.","abstract_has_math":false,"creators":["Mohajer, Mohammad Mahdi"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":["Wang, Song"],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-07-18","date_published":"2024-07-18","updated_at":"2026-07-24T06:33:51Z","subjects":["Computer science","Computer engineering"],"languages":["en"],"rights":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/10315/42143","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Wang, Song"]},{"key":"dc:creator","label":"Author","values":["Mohajer, Mohammad Mahdi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2024-07-18T21:19:53Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2024-07-18T21:19:53Z"]},{"key":"dc:date.issued","label":"Date","values":["2024-07-18"]},{"key":"dc:type","label":"Dc Type","values":["Electronic Thesis or Dissertation"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Computer science","Computer engineering"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10315/42143"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The fairness of machine learning (ML) approaches is critical to the reliability of modern artificial intelligence systems. Despite extensive study on this topic, the fairness of ML models in the software engineering (SE) domain has yet to be well explored. As a result, many ML-powered software systems, particularly those utilized in the software engineering community, continue to be prone to fairness issues. Taking one of the typical SE tasks, i.e., code reviewer recommendation, as a subject, this work conducts the first study toward investigating the fairness of ML applications in the SE domain, explicitly focusing on the code reviewer recommendation task. Our empirical study demonstrates that current state-of-the-art ML-based code reviewer recommendation techniques exhibit unfairness and discriminating behaviors. Specifically, male reviewers get, on average, 7.25% more recommendations than female code reviewers compared to their distribution in the reviewer set. This work also discusses why the studied ML-based code reviewer recommendation systems are unfair and provides solutions to mitigate the unfairness. For instance, these techniques may recommend male reviewers at a significantly higher rate than female reviewers in a discriminatory manner. Our study further indicates that existing mitigation methods can significantly enhance fairness in projects with a similar distribution of protected and privileged groups. Still, their effectiveness in improving fairness on imbalanced or skewed data is limited. Eventually, we suggest a solution to overcome the drawbacks of existing mitigation techniques and tackle bias in imbalanced or skewed datasets."]},{"key":"dc:title","label":"Title","values":["A First Look at Fairness of Machine Learning Based Code Reviewer Recommendation"]}]}],"canonical_facts":{"dc:contributor.advisor":["Wang, Song"],"dc:creator":["Mohajer, Mohammad Mahdi"],"dc:date.accessioned":["2024-07-18T21:19:53Z"],"dc:date.available":["2024-07-18T21:19:53Z"],"dc:date.issued":["2024-07-18"],"dc:description.abstract":["The fairness of machine learning (ML) approaches is critical to the reliability of modern artificial intelligence systems. Despite extensive study on this topic, the fairness of ML models in the software engineering (SE) domain has yet to be well explored. As a result, many ML-powered software systems, particularly those utilized in the software engineering community, continue to be prone to fairness issues. Taking one of the typical SE tasks, i.e., code reviewer recommendation, as a subject, this work conducts the first study toward investigating the fairness of ML applications in the SE domain, explicitly focusing on the code reviewer recommendation task. Our empirical study demonstrates that current state-of-the-art ML-based code reviewer recommendation techniques exhibit unfairness and discriminating behaviors. Specifically, male reviewers get, on average, 7.25% more recommendations than female code reviewers compared to their distribution in the reviewer set. This work also discusses why the studied ML-based code reviewer recommendation systems are unfair and provides solutions to mitigate the unfairness. For instance, these techniques may recommend male reviewers at a significantly higher rate than female reviewers in a discriminatory manner. Our study further indicates that existing mitigation methods can significantly enhance fairness in projects with a similar distribution of protected and privileged groups. Still, their effectiveness in improving fairness on imbalanced or skewed data is limited. Eventually, we suggest a solution to overcome the drawbacks of existing mitigation techniques and tackle bias in imbalanced or skewed datasets."],"dc:identifier.uri":["https://hdl.handle.net/10315/42143"],"dc:language":["en"],"dc:rights":["Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests."],"dc:subject":["Computer science","Computer engineering"],"dc:title":["A First Look at Fairness of Machine Learning Based Code Reviewer Recommendation"],"dc:type":["Electronic Thesis or Dissertation"]},"updated_at":"2026-07-24T06:33:51Z"}