Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 29 for “"N-grams"”.

  1. Part of speech N-grams for information retrieval

    … modelled in the form of part of speech (POS) n-grams, which are contiguous sequences of parts of speech, extracted from text. The distribution of POS n-grams in language is statistically analysed. It is shown that there exists a relationship between the frequency and informative content of POS …

    glasgow Repository record for Part of speech N-grams for information retrieval (opens in a new tab)

  2. NeuroYara: Learning to Rank for Yara Rules Generation through Deep Language Modeling & Discriminative N-gram Encoding

    … and non-inclusive database of hard-coded byte n-grams. This database is used as the reference set for the automated Yara rules generator which aids in reducing the number of false-positive predictions. Hence, instead of storing a huge non-inclusive database to score byte n-grams, we propose a …

    queens Repository record for NeuroYara: Learning to Rank for Yara Rules Generation through Deep Language Modeling & Discriminative N-gram Encoding (opens in a new tab)

  3. Efficient Substring Discovery Using Suffix, LCP Array and Algorithm-Architecture Interaction

    … and LCP array as a perfect tool to compute N-grams (substring) in various dimensions. Since past couple of decades there has been significant research on construction of suffix and LCP array. Comparatively the research of properly utilizing this prospective data structures to retrieve the …

    lsu-thes Repository record for Efficient Substring Discovery Using Suffix, LCP Array and Algorithm-Architecture Interaction (opens in a new tab)

  4. Language Modeling for limited-data domains

    … relevance for each n-gram context. Since the n-grams from partially matched corpora may not be of equal relevance to the target domain, we propose an n-gram weighting scheme to adjust the component n-gram probabilities based on features derived from readily available corpus segmentation and …

    mit Repository record for Language Modeling for limited-data domains (opens in a new tab)

  5. Novel context-based text analytics methods with applications on customer reviews for products and services

    … feedback. The meaning of the words and short n-grams are context dependent, and it must be understood in the sentence that they appear. We use sentences as the basic elements or features for describing text that helps to understand the meaning of words or n-grams in their context. In chapter 2, …

    vt Repository record for Novel context-based text analytics methods with applications on customer reviews for products and services (opens in a new tab)

  6. N-gram models of agreement in language

    … is an exceedingly large number of other n-grams which waste probability mass without making a reciprocal contribution in the formulation of accurate probability estimates. This thesis describes a simple modification to the n-gram approach which attempts to preserve and enhance the most …

    waikato-masters Repository record for N-gram models of agreement in language (opens in a new tab)

  7. Spoken Language Processing and Modeling for Aviation Communications

    … to natural language processing, such as N-grams and word lattices. This thesis experiments with a process for pretraining transformer-based language models on aviation English corpora and compare the effectiveness and performance of language models transfer learned from pretrained …

    embry-riddle Repository record for Spoken Language Processing and Modeling for Aviation Communications (opens in a new tab)

  8. Support vector machines, N-gram kernels, and text classification

    … In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>

    eastern-wash Repository record for Support vector machines, N-gram kernels, and text classification (opens in a new tab)

  9. StoryPass: a system and study for memorable secure passphrases.

    … performed through an algorithm that uses n-grams to estimate the number of attempts required to successfully guess passphrases created in StoryPass. We were able to successfully guess 64% of the passphrases collected during our 39-day user study, but with only a very large number of …

    uoit Repository record for StoryPass: a system and study for memorable secure passphrases. (opens in a new tab)

  10. Music signal processing for automatic extraction of harmonic and rhythmic information

    … layer. We investigate the usage of standard N-grams and Factored Language Models (FLM) for automatic chord recognition. Another central topic of this work is the feature extraction techniques. We develop a set of new features that belong to chroma family. A set of novel chroma features that is …

    trento Repository record for Music signal processing for automatic extraction of harmonic and rhythmic information (opens in a new tab)

  11. Detecting grammatical errors with treebank-induced, probabilistic parsers

    … English LFG, and one based on part-of-speech n-grams. In addition, the baseline methods and the new methods are combined in a machine learning-based framework, yielding further improvements.

    dcu Repository record for Detecting grammatical errors with treebank-induced, probabilistic parsers (opens in a new tab)

  12. Improving Companion AI Behavior in MimicA

    … study found that the our implementation of n-grams was successful and 19 of 26 believed our framework would be useful to a game developer.</p>

    calpoly Repository record for Improving Companion AI Behavior in MimicA (opens in a new tab)

  13. Dealing with linguistic mismatches for automatic speech recognition

    … constraints provided by phone-level n-grams. In order to address the issues of linguistic mismatches for current ASR systems, my dissertation investigates both knowledge-gnostic and knowledge-agnostic solutions. In the first part, classic theories relevant to acoustics and articulatory …

    uiuc Repository record for Dealing with linguistic mismatches for automatic speech recognition (opens in a new tab)

  14. Anomaly detection for HTTP intrusion detection : algorithm comparisons and the effect of generalization on accuracy

    … (deterministic finite automaton induction and n-grams) show more promise. This dissertation shows that accurate anomaly detection requires carefully controlled generalization. Too much or too little will result inaccurate results. Calculating the growth rate of the set that describes the anomaly …

    unm Repository record for Anomaly detection for HTTP intrusion detection : algorithm comparisons and the effect of generalization on accuracy (opens in a new tab)

  15. Beyond topic-based representations for text mining

    … is to segment text into words and record their n-grams. While simple term features perform relatively well in topic-based tasks, not all downstream applications are of a topical nature and can be captured by words alone. For example, determining the native language of an English essay writer will …

    uiuc Repository record for Beyond topic-based representations for text mining (opens in a new tab)

  16. Domain-specific lexicon generation for emotion detection from text.

    … conjunction with other representations such as n-grams, part-of-speech and sentiment information. Thirdly we propose two different methods which jointly use an emotion-labelled corpus of tweets and emotion-sentiment mapping proposed in psychology to learn word-level numerical quantification of …

    rgu Repository record for Domain-specific lexicon generation for emotion detection from text. (opens in a new tab)

  17. Benchmarking authorship attribution techniques using over a thousand books by fifty Victorian era novelists

    … using different features such as bag of words, n-grams or newly developed techniques like Word2Vec. To improve our success rate, we have combined some useful features some of which are diversity measure of text, bag of words, bigrams, specific words that are written differently between English and …

    iupui Repository record for Benchmarking authorship attribution techniques using over a thousand books by fifty Victorian era novelists (opens in a new tab)

  18. Inferring human personality from written media

    … participants in these corpora, I extracted unigrams, bigrams and trigrams (n-grams) of tokens and their POS tags, and counted every word/tag permutation that appeared. I considered only features appearing one or more times per 1000 words in the Forum corpus because there was not enough data to …

    hawaii Repository record for Inferring human personality from written media (opens in a new tab)

Page 1 of 2