Back to results

Universitat Oberta de Catalunya (UOC)

Engineering data-sharing practices for a fair and trustworthy AI

Abstract

dc:description.abstract

Machine learning (ML) technology may discriminate toward specific social groups. For example, recent research have revealed that ML applications are more likely to fail in identifying women than males in hospitals. Recent research has identified the data used to train these models as one of the causes of these issues. The research community has proposed guidelines to detect the dimensions that can generate these discriminatory behaviors. However, these proposals lack a set structure, restricting their computation and the creation of engineering approaches built upon them. This thesis presents a domain-specific language to document data for ML. This language has served as a basis for creating the responsible AI extension of \emph{Croissant}, a standard adopted by major search engines, such as \emph{Google Dataset Search}. Moreover, this thesis studies the use of large language models (LLM) to automatically create data documentation and the readiness of scientific data for its use in ML.

Degree

thesis:*
Grantor dc:publisher
Universitat Oberta de Catalunya (UOC)
Year dc:date.issued
2024

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Giner Miguelez, Joan

Subjects

dc:subject × 14

Rights

dc:rights
Statement dc:rights
  • CC BY-NC-ND
Language dc:language.iso
eng

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10609/151212
OAI identifier oai:identifier
oai:openaccess.uoc.edu:10609/151212

Chain of custody

source
Harvested from
Universitat Oberta de Catalunya
Base URL
openaccess.uoc.edu/server/oai/request
Last updated
2026-07-27
Source record
OAI-PMH GetRecord
citation

Giner Miguelez, Joan. Engineering data-sharing practices for a fair and trustworthy AI. Universitat Oberta de Catalunya (UOC), 2024. https://hdl.handle.net/10609/151212