Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 8 of 8 for “"tokenisation"”.

  1. Subword segmental neural language generation for Nguni languages

    … intrinsic evaluation scores than tokenisation-based language models, on average across the four Nguni languages. We also evaluate SSLM as an unsupervised morphological segmenter, showing that its learned subwords are closer to morphemes than standard subword tokens. Since SSLM is …

    cape-town Repository record for Subword segmental neural language generation for Nguni languages (opens in a new tab)

  2. Investigating the future of digital assets in financial services

    … on their associated process of conversion called Tokenisation. Tokenisation is examined by comparing it to the process of digitalisation and a conceptual framework is derived which can be used to evaluate the suitability of traditional financial assets to be converted to digital assets through …

    cork Repository record for Investigating the future of digital assets in financial services (opens in a new tab)

  3. Active and Semi-Supervised Learning for Speech Recognition

    … thesis proposes two methods to improve the input tokenisation that is used to derive the training targets that are used in masked-prediction pre-training; a form of self-supervised learning. The first method is biased self-supervised learning. Instead of clustering the embeddings of a model …

    cambridge Repository record for Active and Semi-Supervised Learning for Speech Recognition (opens in a new tab)

  4. Phonological Representations in Language Models

    … segmentation methods and recent advances in LLM tokenisation, the thesis proposes a new linguistically-motivated subword tokenisation method that achieves comparable compression to existing techniques while improving morphological alignment. Overall, this work demonstrates that phoneme-based …

    cambridge Repository record for Phonological Representations in Language Models (opens in a new tab)

  5. From data to model behaviour: A causal and empirical analysis of how data shapes the behaviour of language models

    … the Regression Discontinuity Design to measure tokenisation bias---how vocabulary choices causally affect probability assignments. The analysis reveals that character spans represented as single tokens receive up to 17 times more probability than when split, with effects persisting even in …

    cambridge Repository record for From data to model behaviour: A causal and empirical analysis of how data shapes the behaviour of language models (opens in a new tab)

  6. An investigation of the impact of NFTs on the Modern Art Market

    … achieved using NFTs, smart contracts and tokenisation on the blockchain. This thesis lays the foundation for subsequent studies to explore other substantial segments of the NFT market, especially the collectables and gaming segments

    cape-town Repository record for An investigation of the impact of NFTs on the Modern Art Market (opens in a new tab)

  7. Extraction of chemical structures and reactions from the literature

    … OPSIN employs a regular grammar to direct tokenisation and parsing leading to the generation of an XML parse tree. Nomenclature operations are applied successively to the tree with many requiring the manipulation of an in-memory connection table representation of the structure under …

    cambridge Repository record for Extraction of chemical structures and reactions from the literature (opens in a new tab)

  8. From big data to personal narratives: a supervised learning framework for decoding the course of traumatic brain injury in intensive care

    … prediction, expanding the predictor set with the tokenisation-embedding encoder (i.e., making models ‘wider’) significantly improves prediction performance whilst adding hidden layers does not (i.e., making models ‘deeper’). Functional outcome prediction is more difficult at higher GOSE thresholds …

    cambridge Repository record for From big data to personal narratives: a supervised learning framework for decoding the course of traumatic brain injury in intensive care (opens in a new tab)