Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 29 for “"N-grams"”.
-
Part of speech N-grams for information retrieval
… modelled in the form of part of speech (POS) n-grams, which are contiguous sequences of parts of speech, extracted from text. The distribution of POS n-grams in language is statistically analysed. It is shown that there exists a relationship between the frequency and informative content of POS …
-
NeuroYara: Learning to Rank for Yara Rules Generation through Deep Language Modeling & Discriminative N-gram Encoding
… and non-inclusive database of hard-coded byte n-grams. This database is used as the reference set for the automated Yara rules generator which aids in reducing the number of false-positive predictions. Hence, instead of storing a huge non-inclusive database to score byte n-grams, we propose a …
-
Efficient Substring Discovery Using Suffix, LCP Array and Algorithm-Architecture Interaction
… and LCP array as a perfect tool to compute N-grams (substring) in various dimensions. Since past couple of decades there has been significant research on construction of suffix and LCP array. Comparatively the research of properly utilizing this prospective data structures to retrieve the …
-
Language Modeling for limited-data domains
… relevance for each n-gram context. Since the n-grams from partially matched corpora may not be of equal relevance to the target domain, we propose an n-gram weighting scheme to adjust the component n-gram probabilities based on features derived from readily available corpus segmentation and …
-
Novel context-based text analytics methods with applications on customer reviews for products and services
… feedback. The meaning of the words and short n-grams are context dependent, and it must be understood in the sentence that they appear. We use sentences as the basic elements or features for describing text that helps to understand the meaning of words or n-grams in their context. In chapter 2, …
-
N-gram models of agreement in language
… is an exceedingly large number of other n-grams which waste probability mass without making a reciprocal contribution in the formulation of accurate probability estimates. This thesis describes a simple modification to the n-gram approach which attempts to preserve and enhance the most …
-
Spoken Language Processing and Modeling for Aviation Communications
… to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained …
-
Support vector machines, N-gram kernels, and text classification
… In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>
-
StoryPass: a system and study for memorable secure passphrases.
… performed through an algorithm that uses n-grams to estimate the number of attempts required to successfully guess passphrases created in StoryPass. We were able to successfully guess 64% of the passphrases collected during our 39-day user study, but with only a very large number of …
-
Music signal processing for automatic extraction of harmonic and rhythmic information
… layer. We investigate the usage of standard N-grams and Factored Language Models (FLM) for automatic chord recognition. Another central topic of this work is the feature extraction techniques. We develop a set of new features that belong to chroma family. A set of novel chroma features that is …
-
Detecting grammatical errors with treebank-induced, probabilistic parsers
… English LFG, and one based on part-of-speech n-grams. In addition, the baseline methods and the new methods are combined in a machine learning-based framework, yielding further improvements.
-
Improving Companion AI Behavior in MimicA
… study found that the our implementation of n-grams was successful and 19 of 26 believed our framework would be useful to a game developer.</p>
-
Dealing with linguistic mismatches for automatic speech recognition
… constraints provided by phone-level n-grams. In order to address the issues of linguistic mismatches for current ASR systems, my dissertation investigates both knowledge-gnostic and knowledge-agnostic solutions. In the first part, classic theories relevant to acoustics and articulatory …
-
Anomaly detection for HTTP intrusion detection : algorithm comparisons and the effect of generalization on accuracy
… (deterministic finite automaton induction and n-grams) show more promise. This dissertation shows that accurate anomaly detection requires carefully controlled generalization. Too much or too little will result inaccurate results. Calculating the growth rate of the set that describes the anomaly …
-
Beyond topic-based representations for text mining
… is to segment text into words and record their n-grams. While simple term features perform relatively well in topic-based tasks, not all downstream applications are of a topical nature and can be captured by words alone. For example, determining the native language of an English essay writer will …
-
Domain-specific lexicon generation for emotion detection from text.
… conjunction with other representations such as n-grams, part-of-speech and sentiment information. Thirdly we propose two different methods which jointly use an emotion-labelled corpus of tweets and emotion-sentiment mapping proposed in psychology to learn word-level numerical quantification of …
-
Benchmarking authorship attribution techniques using over a thousand books by fifty Victorian era novelists
… using different features such as bag of words, n-grams or newly developed techniques like Word2Vec. To improve our success rate, we have combined some useful features some of which are diversity measure of text, bag of words, bigrams, specific words that are written differently between English and …
-
A Comparative Study of Data Transformations for Efficient XML and JSON Data Compression. An In-Depth Analysis of Data Transformation Techniques, including Tag and Capital Conversions, Character and Word N-Gram Transformations, and Domain-Specific Data Transforms using SMILES Data as a Case Study
XML is a widely used data exchange format. The verbose nature of XML leads to the requirement to efficiently store and process this type of data using compression. Various general-purpose transforms and compression techniques exist that can be used to transform and compress XML data. More compact …
-
Inferring human personality from written media
… participants in these corpora, I extracted unigrams, bigrams and trigrams (n-grams) of tokens and their POS tags, and counted every word/tag permutation that appeared. I considered only features appearing one or more times per 1000 words in the Forum corpus because there was not enough data to …
Page 1 of 2