Back to results

Stellenbosch : Stellenbosch University

A framework for evaluating semi-structured hierarchical data using language models

Abstract

dc:description.abstract

Recent advances in large language models have intensified interest in applying them to tasks such as extractive question answering over semi-structured hierarchical data represented in markup languages. The conventional practice of linearising markup to plain text involves removing structural information that is integral to interpretation, and existing structure-aware approaches are often supported by ad hoc, task-specific pipelines. Limited guidance is available on how heterogeneous sources should be transformed into structure-preserving representations, how alternative model adaptation strategies should be compared, and how the impact of structural information should be quantified in a reproducible manner. This has resulted in fragmented methodologies for the development and assessment of language model-based systems for semi-structured data. A generic framework is proposed in this thesis for the processing and evaluation of semi-structured hierarchical data by language models in the context of extractive question answering. The framework is specified as a modular architecture comprising a data preparation component for transforming heterogeneous markup and tabular sources into a canonical markup-rich representation and, where required, synthesising labelled question answering pairs; a model training component for configuring and adapting pre-trained models; and a performance evaluation component for computing text-based and structure-aware metrics as well as organising structured experimental comparisons. The framework is intended to provide a principled basis on which markup-aware question answering systems may be developed and analysed across application domains. A proof-of-concept instantiation of the framework is implemented and subjected to verification and validation. Verification is conducted by applying the instantiation to a web-based HTML question answering benchmark, confirming that performance comparable with reported baselines is attained and that discarding structural information in favour of text-only input leads to measurable degradation. The practical utility and robustness of the framework are then assessed by carrying out various case studies involving semi-structured tables, combined tabular and textual sources, and synthetic relational data. Across these studies, configurations that exploit markup structure consistently yield higher scores in respect of standard evaluation metrics, thereby supporting the contention that structural information of semi-structured documents constitutes a primary signal for language model-based extractive question answering.

Degree

thesis:*
Grantor dc:publisher
Stellenbosch : Stellenbosch University
Year dc:date.issued
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Du Plessis, Stephan Visser
Advisors dc:contributor.advisor
  • Van Vuuren, J. H.
  • Nel, G. S.

Rights

Language dc:language.iso
en

Identifiers

dc:identifier.*
Repository record dc:identifier.uri
https://scholar.sun.ac.za/handle/10019.1/135772
OAI identifier oai:identifier
oai:scholar.sun.ac.za:10019.1/135772

Chain of custody

source
Harvested from
Stellenbosch University
Base URL
scholar.sun.ac.za/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
related terms
citation

Du Plessis, Stephan Visser. A framework for evaluating semi-structured hierarchical data using language models. Stellenbosch : Stellenbosch University, 2026. https://scholar.sun.ac.za/handle/10019.1/135772