Back to results

Reykjavík University

When more is less : identifying biases in large Icelandic corpora

Abstract

dc:description.abstract

Language is a fundamental part of human communication and as technology advances, people expect to be able to communicate, not only with other people, but also with computer devices, in their own language. Linguistic corpora are used for this purpose, to train language models which can predict and model human languages. Various corpora have been assembled for Icelandic and language models have been trained on them. These models predict words and sentences, and help with translating texts and solving various other tasks. In many cases, these models are expected to be representative of modern Icelandic, but is that truly the case? It is sometimes assumed that having more data is better, but if the data is sampled from the same sources or categories of texts, this sampling bias increases the influence of overrepresented sources - therefore, a larger dataset can lead to more bias. The Icelandic Gigaword Corpus (IGC) contains a disproportionate amount of parliament speeches and news articles, which are male dominated fields. Thus, it can be expected that a gender bias will be reflected in the models trained on the corpus. In this project, I seek to evaluate the representativeness of Icelandic corpora, as well as evaluating sampling- and societal biases that may occur in them and models trained on them. I apply exploratory analysis as well as the Word Embedding Association Test (WEAT) to demonstrate biases in the IGC. Additionally, I present Flóra, a small corpus of texts from a feminist magazine, and compare its vocabulary to that of the IGC. The corpus contains only 229,936 tokens, while the IGC contains over 1.5 billion tokens. I will demonstrate that Flóra contains multiple words that do not appear in the IGC and that collecting texts from diverse sources can thus increase the representativeness of a corpus.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Tinna Þuríður Sigurðardóttir 1990-
Contributors dc:contributor
  • Háskólinn í Reykjavík

Subjects

dc:subject × 9

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
http://hdl.handle.net/1946/39430
OAI identifier oai:identifier
oai:skemman.is:1946/39430

Chain of custody

source
Harvested from
Reykjavík University
Base URL
skemman.is/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Tinna Þuríður Sigurðardóttir 1990-. When more is less : identifying biases in large Icelandic corpora. 2021. http://hdl.handle.net/1946/39430