Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 46 for “"N-gram"”.
-
N-gram models of agreement in language
Conventional n-gram language models are well-established as powerful yet simple mechanisms for characterising language structure when low data complexity is the primary objective. Much of their predictive power can be traced to a relatively small number of common word sequences usually comprised of …
-
Support vector machines, N-gram kernels, and text classification
… In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>
-
NeuroYara: Learning to Rank for Yara Rules Generation through Deep Language Modeling & Discriminative N-gram Encoding
… and non-inclusive database of hard-coded byte n-grams. This database is used as the reference set for the automated Yara rules generator which aids in reducing the number of false-positive predictions. Hence, instead of storing a huge non-inclusive database to score byte n-grams, we propose a …
-
Using a symbolic language parser to Improve Markov language models
… natural language processing that combines an n-gram (Markov) model with a symbolic parser. In concert these two techniques are applied to the problem of sentence simplification. The n-gram system is comprised of a relational database backend with a frontend application that presents a …
-
Language Modeling for limited-data domains
… component confidence and relevance for each n-gram context. Since the n-grams from partially matched corpora may not be of equal relevance to the target domain, we propose an n-gram weighting scheme to adjust the component n-gram probabilities based on features derived from readily available …
-
Transformations for linguistic steganography
… substitution checkers are based on contextual n-gram counts and the αskew divergence of those counts derived from the Google n-gram corpus. For adjective deletion, we propose an n-gram count method similar to the substitution n-gram checker and a support vector machine classifier using n-gram …
-
A Comparative Study of Data Transformations for Efficient XML and JSON Data Compression. An In-Depth Analysis of Data Transformation Techniques, including Tag and Capital Conversions, Character and Word N-Gram Transformations, and Domain-Specific Data Transforms using SMILES Data as a Case Study
XML is a widely used data exchange format. The verbose nature of XML leads to the requirement to efficiently store and process this type of data using compression. Various general-purpose transforms and compression techniques exist that can be used to transform and compress XML data. More compact …
-
The predictability problem
… Kombination objektiver Maße (semantische und n-gram-Maße) geschätzt werden kann, die auf den statistischen Eigenschaften von Textkorpora beruhen. Die semantischen Maße werden entweder durch Abfragen von Internet-Suchmaschinen oder durch die Anwendung der Latent Semantic Analysis gebildet, …
-
Ordering prenominal modifiers with a ranking approach
… a strong baseline that makes use of the Google n-gram corpus. We attain a maximum error reduction of 69.8% and average error reduction across all test sets of 59.1% compared to the state-of-the-art, and we attain a maximum error reduction of 68.4% and average error reduction across all test sets …
-
Treebank-based acquisition of Chinese LFG resources for parsing and generation
… acquire robust,wide-coverage Lexical-Functional Grammar (LFG) resources for Chinese parsing and generation, which is part of a larger project on the rapid construction of deep, large-scale, constraint-based, multilingual grammatical resources. I present an application-oriented LFG analysis for …
-
Direct multidisplay for web document repositories
… intensive and time consuming. Multibrowser, a program that addresses this problem, is presented in this thesis. Multibrowser combines the advantages of multidisplay and direct display to present a more efficient user computer interface. First, the system downloads the actual documents according to …
-
Inference-Time Learning Algorithms of Language Models
… circuits that implement approximate n-gram learning algorithms for probabilistic languages. Building on these insights, I develop two approaches to enhance LMs. First, I demonstrate that explicitly incorporating n-gram computation into model architectures improves performance across …
-
Surface Realization Using a Featurized Syntactic Statistical Language Model
… surface realization is the generation of grammatical sentences from incomplete sentence plans. Realization can be broken into a two-stage process consisting of an over-generating rule-based module followed by a ranker that outputs the most probable candidate sentence based on a statistical …
-
Query Segmentation For E-Commerce Sites
… focus on the web search using Google n-gram frequencies corpus or text retrieval from relational databases. However, this module is also useful in the domain of E-Commerce for product search. In this thesis, we will discuss query segmentation in the context of the E-Commerce area. We …
-
Photo annotation and retrieval through speech
… speech-based retrieval we have developed a mixed grammar recognition approach which allows the speech recognition system to construct a single finite-state network combining context-free grammars for recognizing and parsing query carrier phrases and metadata phrases, with an unconstrained …
-
Automatic utterance segmentation in spontaneous speech
… the incorporation of pause duration and grammar information to the utterance segmentation task. As a result, we obtain an optimal set of parameters for the lower level utterance segmenter, and show that part-of-speech based N-gram language modeling of the spoken words in conjunction with …
-
Inferring insulin regimen from clinical notes : using natural language processing techniques to extract data from free text records
… in outpatient clinical notes. We explore two n-gram models - Logistic Regression and Conditional Random Field and analyze their performance. We also explore models using contextual word representations from the domain specific pretrained language models, character level embeddings and auxillary …
-
Efficient Substring Discovery Using Suffix, LCP Array and Algorithm-Architecture Interaction
… and LCP array as a perfect tool to compute N-grams (substring) in various dimensions. Since past couple of decades there has been significant research on construction of suffix and LCP array. Comparatively the research of properly utilizing this prospective data structures to retrieve the …
-
There Is Always an Option
… I develop interpretation methods based on $n$-gram statistics of action sequences and mean+variance spatial mappings of agent states. These analyses show that the clusters match clear behaviours and movement patterns when the agent follows an option. Overall, the tool enables the detection of …
-
Evaluation of Automatic Text Summarization Using Synthetic Facts
… correlation with human judgment with existing N-gram overlap-based metrics such as ROUGE and BLEU and a BERT-based evaluation metric, BERTScore. Our system's experimental evaluation of PEGASUS, BART, and T5 outperforms the current evaluation metrics in measuring factual consistency with a …
Page 1 of 3