Back to results

Oxford Brookes University

Occupational gender bias in large language models : a multi-level analysis using the OccuBias Dataset

Abstract

dc:description

Large language models (LLMs) are increasingly integrated into systems that influence employment pathways, including AI-powered résumé screening platforms (e.g., HireVue and LinkedIn Recruiter) and interview preparation tools such as Google’s Interview Warmup. Although these systems are designed to improve efficiency and scalability, they are often trained on large corpora of historical text that reflect existing social patterns in the labor market. As a result, they may unintentionally reproduce or reinforce occupational gender stereotypes. This study investigates occupational gender bias in two open-source LLMs, Meta’s LLaMA-3-8B-Instruct and Google’s Gemma-7B-IT, using a purpose-built evaluation dataset named OccuBias, comprising 6,300 prompts across gendered, gender-neutral, and stereotype-challenging occupational contexts. Bias is analyzed at both the output level and the embedding level using three complementary metrics: cosine similarity, Scoring Association Means of Word Embeddings (SAME), and direct bias scores. In addition, K-means clustering is applied to explore how occupations are organized within the models’ embedding spaces, and correlation analysis is used to examine the relationship between embedding-level bias and observable gender patterns in generated outputs. The results reveal clear differences between the two models. Gemma-7B-IT demonstrates stronger male-associated tendencies, producing more than 76% male pronoun assignments under deterministic decoding, whereas LLaMA-3-8B-Instruct produces over 63% male pronoun assignments, indicating a comparatively weaker but still noticeable male skew. At the embedding level, Gemma-7B-IT exhibits a stronger alignment between representational bias and generated outputs, while LLaMA-3-8B-Instruct displays more diffuse and less consistent bias patterns. Across both models, the Scoring Association Means of Word Embeddings (SAME) metric proves more sensitive than cosine similarity in detecting subtle gender associations within the embedding space. This research makes four contributions: (i) the introduction of OccuBias, an openly available occupational gender bias dataset designed for systematic evaluation of large language models; (ii) a reproducible multi-level evaluation framework that integrates prompt-based behavioral analysis with embedding-level bias measurement; (iii) comparative insights into gender bias across two contemporary instruction-tuned LLMs, LLaMA-3-8B-Instruct and Gemma-7B-IT; and (iv) an analysis of the relationship between embedding-level bias and output-level gender predictions. These findings may offer practical value for researchers, developers, and policymakers working to build fairer and more transparent AI systems across domains where automated language processing shapes human opportunities, identities, and experiences.

Degree

thesis:*
Grantor dc:publisher
Oxford Brookes University

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zaman, Nafisa Salsabil
Contributors dc:contributor
  • Rolf, Matthias
  • Crook, Nigel

Rights

dc:rights
Statement dc:rights
  • All rights reserved
Language dc:language
en

Identifiers

dc:identifier.*
OAI identifier oai:identifier
tle:d5f1421e-9a5f-4542-83ba-d7e6c643c518:d6bd9758-527a-46cd-bfe2-c433766e8fca:1

Chain of custody

source
Harvested from
Oxford Brookes University
Base URL
radar.brookes.ac.uk/radar/oai
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Zaman, Nafisa Salsabil. Occupational gender bias in large language models : a multi-level analysis using the OccuBias Dataset. Oxford Brookes University, https://doi.org/10.24384/fcfy-3193