Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 26 for “"part-of-speech tagging"”.
-
Speaker Dependent Voice Recognition with Word-Tense Association and Part-of-Speech Tagging
<p>Extensive Research has been conducted on speech recognition and Speaker Recognition over the past few decades. Speaker recognition deals with identifying the speaker from multiple speakers and the ability to filter out the voice of an individual from the background for computational …
-
Development of part of speech tagging and syntactic analysis software for Chinese text
Thesis (M.Eng. and S.B.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, June 2001.
-
Automatic Lexicon Generation for Unsupervised Part-of-Speech Tagging Using Only Unannotated Text
With the growing number of textual resources available, the ability to understand them becomes critical. An essential first step in understanding these sources is the ability to identify the parts-of-speech in each sentence. The goal of this research is to propose, improve, and implement an …
-
The effects of part–of–speech tagging on text–to–speech synthesis for resource–scarce languages
In the world of human language technology, resource-scarce languages (RSLs) suffer from the problem of little available electronic data and linguistic expertise. The Lwazi project in South Africa is a large-scale endeavour to collect and apply such resources for all eleven of the official South …
-
Part-of-speech tagging and partial parsing for Irish using finite-state transducers and constraint grammar
… we present the development and evaluation of a suite of annotation tools for unrestricted Irish text, which go from tokenization, morphological analysis, part-of-speech tagging, right through to partial parsing. In order to develop such tools, a large body of texts is required for testing …
-
Multi-Class Classification in Natural Language Processing
… and empirical arguments for the advantages of using: (i) Sentence structure. (ii) The Sequential Model . Empirical arguments are given using word-prediction and part of speech tagging tasks. Theoretical arguments present this thesis as an extension of the current classification methods which …
-
Machine Learning for Information Extraction
The dissertation presents a number of novel machine learning techniques and applies them to information extraction. The study addresses several information extraction subtasks: part of speech tagging, entity extraction, coreference resolution, and relation extraction. Each of the tasks is …
-
Multi-source domain adaptation with mixture of experts
We propose a mixture-of-experts approach for unsupervised domain adaptation from multiple sources. The key idea is to explicitly capture the relationship between a target example and different source domains. This relationship, expressed by a point-to-set metric, determines how to combine …
-
Unsupervised multilingual learning
… languages. In this thesis, we present a class of probabilistic models that exploit these links as a form of naturally occurring supervision. These models allow us to substantially improve performance for core text processing tasks, such as morphological segmentation, part-of-speech tagging, and …
-
The Application of P-Bar Theory in Transformation-Based Error-Driven Learning
… a rule based method for determining the context of a partext (i.e., a part of a text document). </p> <p>In <em>Transformation-Based Error-Driven Learning and Natural Language Processing: A Case Study in Part-of-Speech Tagging </em>Brill (1995) demonstrates a method of error-driven learning …
-
Efficient Lagrangian relaxation algorithms for exact inference in natural language tasks
… best solution requires a search over a large set of possible structures. Solving these combinatorial search problems exactly can be inefficient, and so researchers often use approximate techniques at the cost of model accuracy. In this thesis, we turn to Lagrangian relaxation as an alternative to …
-
Modeling second language learners' interlanguage and its variability: a computer-based dynamic assessment approach to distinguishing between errors and mistakes
… errors and mistakes in texts written by learners of French, to then investigate the extent to which interlanguage competence varies across time, text types, and students. The key outcomes include: 1. An expanded model based on dynamic assessment principles to distinguish between errors and …
-
A log-linear discriminative modeling framework for speech recognition
Conventional speech recognition systems are based on Gaussian hidden Markov models (HMMs).Discriminative techniques such as log-linear modeling have been investigated in speech recognition only recently. This thesis establishes a log-linear modeling framework in the context of discriminative …
-
The computational analysis of morphosyntactic categories in Urdu
Urdu is a language of the Indo-Aryan family, widely spoken in India and Pakistan, and an important minority language in Europe, North America, and elsewhere. This thesis describes the development of a computer-based system for part-of-speech tagging of Urdu texts, consisting of a tagset, a set of …
-
Lagrangian relaxation for natural language decoding
The major success story of natural language processing over the last decade has been the development of high-accuracy statistical methods for a wide-range of language applications. The availability of large textual data sets has made it possible to employ increasingly sophisticated statistical …
-
Non-parametric Bayesian models for structured output prediction
… labels. This means that the presence or value of a given label affects the other labels, for instance in text labelling problems, where output labels are applied to each word, and their interdependencies must be modelled. Non-parametric Bayesian (NPB) techniques are probabilistic modelling …
-
Splitting rocks: Learning word sense representations from corpora and lexica
The representation of written language semantics is a central problem of language technology and a crucial component of many natural language processing applications, from part-of-speech tagging to text summarization. These representations of linguistic units, such as words or sentences, allow …
-
Inductive Bias and Modular Design for Sample-Efficient Neural Language Learning
Most of the world's languages suffer from the paucity of annotated data. This curbs the effectiveness of supervised learning, the most widespread approach to modelling language. Instead, an alternative paradigm could take inspiration from the propensity of children to acquire language from limited …
-
Classification and visualisation of text documents using networks
In both the areas of text classification and text visualisation graph/network theoretic methods can be applied effectively. For text classification we assessed the effectiveness of graph/network summary statistics to develop weighting schemes and features to improve test accuracy. For text …
-
Heuristisk analys med Diderichsens satsschema. Tillämpningar för svensk text
… on main clause (primary) analysis, a collection of licensing techniques for removing non-primary verb candidates is employed, leaving e.g. the primary verbs, particles and conjunctions (bounded key constituents) that delimit the content of the fields in Diderichsen’s sentence schema. Hereby, the …
Page 1 of 2