Stellenbosch : Stellenbosch University
A framework for evaluating semi-structured hierarchical data using language models
Abstract
dc:description.abstractRecent advances in large language models have intensified interest in applying them to tasks such as extractive question answering over semi-structured hierarchical data represented in markup languages. The conventional practice of linearising markup to plain text involves removing structural information that is integral to interpretation, and existing structure-aware approaches are often supported by ad hoc, task-specific pipelines. Limited guidance is available on how heterogeneous sources should be transformed into structure-preserving representations, how alternative model adaptation strategies should be compared, and how the impact of structural information should be quantified in a reproducible manner. This has resulted in fragmented methodologies for the development and assessment of language model-based systems for semi-structured data. A generic framework is proposed in this thesis for the processing and evaluation of semi-structured hierarchical data by language models in the context of extractive question answering. The framework is specified as a modular architecture comprising a data preparation component for transforming heterogeneous markup and tabular sources into a canonical markup-rich representation and, where required, synthesising labelled question answering pairs; a model training component for configuring and adapting pre-trained models; and a performance evaluation component for computing text-based and structure-aware metrics as well as organising structured experimental comparisons. The framework is intended to provide a principled basis on which markup-aware question answering systems may be developed and analysed across application domains. A proof-of-concept instantiation of the framework is implemented and subjected to verification and validation. Verification is conducted by applying the instantiation to a web-based HTML question answering benchmark, confirming that performance comparable with reported baselines is attained and that discarding structural information in favour of text-only input leads to measurable degradation. The practical utility and robustness of the framework are then assessed by carrying out various case studies involving semi-structured tables, combined tabular and textual sources, and synthetic relational data. Across these studies, configurations that exploit markup structure consistently yield higher scores in respect of standard evaluation metrics, thereby supporting the contention that structural information of semi-structured documents constitutes a primary signal for language model-based extractive question answering.
Degree
thesis:*- Grantor dc:publisher
- Stellenbosch : Stellenbosch University
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Du Plessis, Stephan Visser
- Advisors dc:contributor.advisor
-
- Van Vuuren, J. H.
- Nel, G. S.
Rights
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Repository record dc:identifier.uri
- https://scholar.sun.ac.za/handle/10019.1/135772
- OAI identifier oai:identifier
- oai:scholar.sun.ac.za:10019.1/135772