Back to results

University of Illinois at Urbana-Champaign

Exposing and correcting the gender bias in image captioning datasets and models

Abstract

dc:description

The task of image captioning implicitly involves gender identification. However, due to the gender bias in data, gender identification by an image captioning model suffers. Also, due to the word-by-word prediction, the gender-activity bias in the data tends to influence the other words in the caption, resulting in the well know problem of label bias. In this work, we investigate gender bias in the COCO captioning dataset, and show that it engenders not only from the statistical distribution of genders with contexts but also from the flawed per instance annotation provided by the human annotators. We then look at the issues created by this bias in the models trained on the data. We propose a technique to get rid of the bias by splitting the task into 2 subtasks: gender-neutral image captioning and gender classification. By this decoupling, the gender-context influence can be eradicated. We train a gender neutral image captioning model, which does not exhibit the language model based bias arising from the gender and gives good quality captions. This model gives comparable results to a gendered model even when evaluating against a dataset that possesses similar bias as the training data. Interestingly, the predictions by this model on images without humans, are also visibly different from the one trained on gendered captions. For injecting gender into the captions, we train gender classifiers using cropped portions of images that contain only the person. This allows us to get rid of the context and focus on the person to predict the gender. We train bounding box based and body mask based classifiers, giving a much higher accuracy in gender prediction than an image captioning model implicitly attempting to classify the gender from the full image. By substituting the genders into the gender-neutral captions, we get the final gendered predictions. Our predictions achieve similar performance to a model trained with gender, and at the same time are devoid of gender bias. Finally, our main result is that on an anti-stereotypical dataset, our model outperforms a popular image captioning model which is trained with gender.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2019

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Bhargava, Shruti
Contributors dc:contributor
  • Forsyth, David Alexander

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright 2019 Shruti Bhargava
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/105104
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/105104

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Bhargava, Shruti. Exposing and correcting the gender bias in image captioning datasets and models. Thesis thesis, University of Illinois at Urbana-Champaign, 2019. http://hdl.handle.net/2142/105104