{"id":{"repo_id":"reykjavik","oai_identifier":"oai:skemman.is:1946/39430"},"canonical_url":"https://search.dev.ndltd.org/etd/reykjavik/oai:skemman.is:1946/39430","repository":{"repo_id":"reykjavik","name":"Reykjavík University","base_url":"https://skemman.is/oai/request"},"display":{"title":"When more is less : identifying biases in large Icelandic corpora","abstract":"Language is a fundamental part of human communication and as technology advances, people expect to be able to communicate, not only with other people, but also with computer devices, in their own language. Linguistic corpora are used for this purpose, to train language models which can predict and model human languages. Various corpora have been assembled for Icelandic and language models have been trained on them. These models predict words and sentences, and help with translating texts and solving various other tasks. In many cases, these models are expected to be representative of modern Icelandic, but is that truly the case? It is sometimes assumed that having more data is better, but if the data is sampled from the same sources or categories of texts, this sampling bias increases the influence of overrepresented sources - therefore, a larger dataset can lead to more bias. The Icelandic Gigaword Corpus (IGC) contains a disproportionate amount of parliament speeches and news articles, which are male dominated fields. Thus, it can be expected that a gender bias will be reflected in the models trained on the corpus. In this project, I seek to evaluate the representativeness of Icelandic corpora, as well as evaluating sampling- and societal biases that may occur in them and models trained on them. I apply exploratory analysis as well as the Word Embedding Association Test (WEAT) to demonstrate biases in the IGC. Additionally, I present Flóra, a small corpus of texts from a feminist magazine, and compare its vocabulary to that of the IGC. The corpus contains only 229,936 tokens, while the IGC contains over 1.5 billion tokens. I will demonstrate that Flóra contains multiple words that do not appear in the IGC and that collecting texts from diverse sources can thus increase the representativeness of a corpus.","abstract_html":"Language is a fundamental part of human communication and as technology advances, people expect to be able to communicate, not only with other people, but also with computer devices, in their own language. Linguistic corpora are used for this purpose, to train language models which can predict and model human languages. Various corpora have been assembled for Icelandic and language models have been trained on them. These models predict words and sentences, and help with translating texts and solving various other tasks. In many cases, these models are expected to be representative of modern Icelandic, but is that truly the case? It is sometimes assumed that having more data is better, but if the data is sampled from the same sources or categories of texts, this sampling bias increases the influence of overrepresented sources - therefore, a larger dataset can lead to more bias. The Icelandic Gigaword Corpus (IGC) contains a disproportionate amount of parliament speeches and news articles, which are male dominated fields. Thus, it can be expected that a gender bias will be reflected in the models trained on the corpus. In this project, I seek to evaluate the representativeness of Icelandic corpora, as well as evaluating sampling- and societal biases that may occur in them and models trained on them. I apply exploratory analysis as well as the Word Embedding Association Test (WEAT) to demonstrate biases in the IGC. Additionally, I present Flóra, a small corpus of texts from a feminist magazine, and compare its vocabulary to that of the IGC. The corpus contains only 229,936 tokens, while the IGC contains over 1.5 billion tokens. I will demonstrate that Flóra contains multiple words that do not appear in the IGC and that collecting texts from diverse sources can thus increase the representativeness of a corpus.","abstract_has_math":false,"creators":["Tinna Þuríður Sigurðardóttir 1990-"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Háskólinn í Reykjavík"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2021,"date_issued":"2021-06-23T12:36:07Z","date_published":"2021-06-23T12:36:07Z","updated_at":"2026-07-27T20:36:46Z","subjects":["Tölvunarfræði","Meistaraprófsritgerðir","Máltækni","Málskynjun (tölvur)","Reiknilíkön","Computer science","Computational linguistics","Automatic speech recognition","Computer algorithms"],"languages":["en"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1946/39430","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Háskólinn í Reykjavík"]},{"key":"dc:creator","label":"Author","values":["Tinna Þuríður Sigurðardóttir 1990-"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2021-06-23T12:36:03Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2021-06-23T12:36:03Z"]},{"key":"dc:date.issued","label":"Date","values":["2021-06-23T12:36:07Z"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Tölvunarfræði","Meistaraprófsritgerðir","Máltækni","Málskynjun (tölvur)","Reiknilíkön","Computer science","Computational linguistics","Automatic speech recognition","Computer algorithms"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1946/39430"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Language is a fundamental part of human communication and as technology advances, people expect to be able to communicate, not only with other people, but also with computer devices, in their own language. Linguistic corpora are used for this purpose, to train language models which can predict and model human languages. Various corpora have been assembled for Icelandic and language models have been trained on them. These models predict words and sentences, and help with translating texts and solving various other tasks. In many cases, these models are expected to be representative of modern Icelandic, but is that truly the case? It is sometimes assumed that having more data is better, but if the data is sampled from the same sources or categories of texts, this sampling bias increases the influence of overrepresented sources - therefore, a larger dataset can lead to more bias. The Icelandic Gigaword Corpus (IGC) contains a disproportionate amount of parliament speeches and news articles, which are male dominated fields. Thus, it can be expected that a gender bias will be reflected in the models trained on the corpus. In this project, I seek to evaluate the representativeness of Icelandic corpora, as well as evaluating sampling- and societal biases that may occur in them and models trained on them. I apply exploratory analysis as well as the Word Embedding Association Test (WEAT) to demonstrate biases in the IGC. Additionally, I present Flóra, a small corpus of texts from a feminist magazine, and compare its vocabulary to that of the IGC. The corpus contains only 229,936 tokens, while the IGC contains over 1.5 billion tokens. I will demonstrate that Flóra contains multiple words that do not appear in the IGC and that collecting texts from diverse sources can thus increase the representativeness of a corpus.","Tungumál eru mikilvægur hluti af mannlegum samskiptum. Eftir því sem tækninni hefur fleygt fram, höfum við farið að gera meiri kröfur til þess að geta átt samskipti við tæki á okkar eigin tungu. Svokallaðar málheildir, eða málleg gagnasöfn, eru notuð til þess að þjálfa mállíkön sem hafa forspárgildi og geta upp að ákveðnu marki skilið mannamál. Það eru til ýmsar íslenskar málheildir og þær hafa verið notaðar til þess að þjálfa slík mállíkön. Þessi líkön spá fyrir um orð og setningar og hjálpa til við að leysa ýmis verkefni eins og til dæmis þýðingaverkefni eða að búa til spjallmenni. Í mörgum tilvikum er gert ráð fyrir því að þessi líkön nái utanum íslenskt nútímamál, en hvernig getum við fullvissað okkur um það? Við gerum oft ráð fyrir að meira magn sé betra en minna þegar það kemur að gögnum, en ef gögnum er safnað frá hentugleikaúrtaki, getur aukið gagnamagn leitt til aukinnar skekkju. Til að mynda er meira en helmingur Risamálheildar fenginn úr fréttatextum og alþingisræðum, en hallað hefur á konur bæði í íslenskum fjölmiðlum, sem og á Alþingi. Við getum því búist við því að finna kynjahalla í málheildinni og líkönum sem byggja á henni. Í þessu verkefni leitast ég við að meta íslenskar málheildir með tilliti til þess hversu vel þær ná utanum íslenska tungu. Einnig mun ég meta skekkjur tengdar gagnasöfnun og félagslegum þáttum. Í þessum tilgangi mun ég nota eigin þýðingu á Word Embedding Association Test (WEAT), sem hefur verið notað til þess að meta skekkjur í mállíkönum, bæði á ensku og fleiri tungumálum. Ég mun einnig kynna nýja íslenska málheild, Flóru, sem ég hef sett saman til þess að undirstrika hversu auðvelt það er að auka fjölbreytileika stærri málheilda. Þessi málheild er einungis 229.936 orð, en Risamálheildin inniheldur meira en einn og hálfan milljarð orða. Ég mun sýna fram á að Flóra inniheldur fjöldamörg orð sem ekki koma fram í Risamáheild og því er, með nokkuð auðveldum hætti, hægt að ná betur utanum íslenska tungu með því að safna gögnum frá fleiri heimildum."]},{"key":"dc:title","label":"Title","values":["When more is less : identifying biases in large Icelandic corpora"]}]}],"canonical_facts":{"dc:contributor":["Háskólinn í Reykjavík"],"dc:creator":["Tinna Þuríður Sigurðardóttir 1990-"],"dc:date.accessioned":["2021-06-23T12:36:03Z"],"dc:date.available":["2021-06-23T12:36:03Z"],"dc:date.issued":["2021-06-23T12:36:07Z"],"dc:description.abstract":["Language is a fundamental part of human communication and as technology advances, people expect to be able to communicate, not only with other people, but also with computer devices, in their own language. Linguistic corpora are used for this purpose, to train language models which can predict and model human languages. Various corpora have been assembled for Icelandic and language models have been trained on them. These models predict words and sentences, and help with translating texts and solving various other tasks. In many cases, these models are expected to be representative of modern Icelandic, but is that truly the case? It is sometimes assumed that having more data is better, but if the data is sampled from the same sources or categories of texts, this sampling bias increases the influence of overrepresented sources - therefore, a larger dataset can lead to more bias. The Icelandic Gigaword Corpus (IGC) contains a disproportionate amount of parliament speeches and news articles, which are male dominated fields. Thus, it can be expected that a gender bias will be reflected in the models trained on the corpus. In this project, I seek to evaluate the representativeness of Icelandic corpora, as well as evaluating sampling- and societal biases that may occur in them and models trained on them. I apply exploratory analysis as well as the Word Embedding Association Test (WEAT) to demonstrate biases in the IGC. Additionally, I present Flóra, a small corpus of texts from a feminist magazine, and compare its vocabulary to that of the IGC. The corpus contains only 229,936 tokens, while the IGC contains over 1.5 billion tokens. I will demonstrate that Flóra contains multiple words that do not appear in the IGC and that collecting texts from diverse sources can thus increase the representativeness of a corpus.","Tungumál eru mikilvægur hluti af mannlegum samskiptum. Eftir því sem tækninni hefur fleygt fram, höfum við farið að gera meiri kröfur til þess að geta átt samskipti við tæki á okkar eigin tungu. Svokallaðar málheildir, eða málleg gagnasöfn, eru notuð til þess að þjálfa mállíkön sem hafa forspárgildi og geta upp að ákveðnu marki skilið mannamál. Það eru til ýmsar íslenskar málheildir og þær hafa verið notaðar til þess að þjálfa slík mállíkön. Þessi líkön spá fyrir um orð og setningar og hjálpa til við að leysa ýmis verkefni eins og til dæmis þýðingaverkefni eða að búa til spjallmenni. Í mörgum tilvikum er gert ráð fyrir því að þessi líkön nái utanum íslenskt nútímamál, en hvernig getum við fullvissað okkur um það? Við gerum oft ráð fyrir að meira magn sé betra en minna þegar það kemur að gögnum, en ef gögnum er safnað frá hentugleikaúrtaki, getur aukið gagnamagn leitt til aukinnar skekkju. Til að mynda er meira en helmingur Risamálheildar fenginn úr fréttatextum og alþingisræðum, en hallað hefur á konur bæði í íslenskum fjölmiðlum, sem og á Alþingi. Við getum því búist við því að finna kynjahalla í málheildinni og líkönum sem byggja á henni. Í þessu verkefni leitast ég við að meta íslenskar málheildir með tilliti til þess hversu vel þær ná utanum íslenska tungu. Einnig mun ég meta skekkjur tengdar gagnasöfnun og félagslegum þáttum. Í þessum tilgangi mun ég nota eigin þýðingu á Word Embedding Association Test (WEAT), sem hefur verið notað til þess að meta skekkjur í mállíkönum, bæði á ensku og fleiri tungumálum. Ég mun einnig kynna nýja íslenska málheild, Flóru, sem ég hef sett saman til þess að undirstrika hversu auðvelt það er að auka fjölbreytileika stærri málheilda. Þessi málheild er einungis 229.936 orð, en Risamálheildin inniheldur meira en einn og hálfan milljarð orða. Ég mun sýna fram á að Flóra inniheldur fjöldamörg orð sem ekki koma fram í Risamáheild og því er, með nokkuð auðveldum hætti, hægt að ná betur utanum íslenska tungu með því að safna gögnum frá fleiri heimildum."],"dc:identifier.uri":["http://hdl.handle.net/1946/39430"],"dc:language.iso":["en"],"dc:subject":["Tölvunarfræði","Meistaraprófsritgerðir","Máltækni","Málskynjun (tölvur)","Reiknilíkön","Computer science","Computational linguistics","Automatic speech recognition","Computer algorithms"],"dc:title":["When more is less : identifying biases in large Icelandic corpora"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T20:36:46Z"}