Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 46 for “"N-gram"”.

  1. N-gram models of agreement in language

    Conventional n-gram language models are well-established as powerful yet simple mechanisms for characterising language structure when low data complexity is the primary objective. Much of their predictive power can be traced to a relatively small number of common word sequences usually comprised of …

    waikato-masters Repository record for N-gram models of agreement in language (opens in a new tab)

  2. Support vector machines, N-gram kernels, and text classification

    … In this paper we explore kernels based off of N-grams or consecutive sequences of words.</p>

    eastern-wash Repository record for Support vector machines, N-gram kernels, and text classification (opens in a new tab)

  3. NeuroYara: Learning to Rank for Yara Rules Generation through Deep Language Modeling & Discriminative N-gram Encoding

    … and non-inclusive database of hard-coded byte n-grams. This database is used as the reference set for the automated Yara rules generator which aids in reducing the number of false-positive predictions. Hence, instead of storing a huge non-inclusive database to score byte n-grams, we propose a …

    queens Repository record for NeuroYara: Learning to Rank for Yara Rules Generation through Deep Language Modeling & Discriminative N-gram Encoding (opens in a new tab)

  4. Using a symbolic language parser to Improve Markov language models

    … natural language processing that combines an n-gram (Markov) model with a symbolic parser. In concert these two techniques are applied to the problem of sentence simplification. The n-gram system is comprised of a relational database backend with a frontend application that presents a …

    mit Repository record for Using a symbolic language parser to Improve Markov language models (opens in a new tab)

  5. Language Modeling for limited-data domains

    … component confidence and relevance for each n-gram context. Since the n-grams from partially matched corpora may not be of equal relevance to the target domain, we propose an n-gram weighting scheme to adjust the component n-gram probabilities based on features derived from readily available …

    mit Repository record for Language Modeling for limited-data domains (opens in a new tab)

  6. Transformations for linguistic steganography

    … substitution checkers are based on contextual n-gram counts and the αskew divergence of those counts derived from the Google n-gram corpus. For adjective deletion, we propose an n-gram count method similar to the substitution n-gram checker and a support vector machine classifier using n-gram

    cambridge Repository record for Transformations for linguistic steganography (opens in a new tab)

  7. The predictability problem

    … Kombination objektiver Maße (semantische und n-gram-Maße) geschätzt werden kann, die auf den statistischen Eigenschaften von Textkorpora beruhen. Die semantischen Maße werden entweder durch Abfragen von Internet-Suchmaschinen oder durch die Anwendung der Latent Semantic Analysis gebildet, …

    potsdam-diss Repository record for The predictability problem (opens in a new tab)

  8. Ordering prenominal modifiers with a ranking approach

    … a strong baseline that makes use of the Google n-gram corpus. We attain a maximum error reduction of 69.8% and average error reduction across all test sets of 59.1% compared to the state-of-the-art, and we attain a maximum error reduction of 68.4% and average error reduction across all test sets …

    mit Repository record for Ordering prenominal modifiers with a ranking approach (opens in a new tab)

  9. Treebank-based acquisition of Chinese LFG resources for parsing and generation

    … acquire robust,wide-coverage Lexical-Functional Grammar (LFG) resources for Chinese parsing and generation, which is part of a larger project on the rapid construction of deep, large-scale, constraint-based, multilingual grammatical resources. I present an application-oriented LFG analysis for …

    dcu Repository record for Treebank-based acquisition of Chinese LFG resources for parsing and generation (opens in a new tab)

  10. Direct multidisplay for web document repositories

    … intensive and time consuming. Multibrowser, a program that addresses this problem, is presented in this thesis. Multibrowser combines the advantages of multidisplay and direct display to present a more efficient user computer interface. First, the system downloads the actual documents according to …

    iastate Repository record for Direct multidisplay for web document repositories (opens in a new tab)

  11. Inference-Time Learning Algorithms of Language Models

    … circuits that implement approximate n-gram learning algorithms for probabilistic languages. Building on these insights, I develop two approaches to enhance LMs. First, I demonstrate that explicitly incorporating n-gram computation into model architectures improves performance across …

    mit Repository record for Inference-Time Learning Algorithms of Language Models (opens in a new tab)

  12. Surface Realization Using a Featurized Syntactic Statistical Language Model

    … surface realization is the generation of grammatical sentences from incomplete sentence plans. Realization can be broken into a two-stage process consisting of an over-generating rule-based module followed by a ranker that outputs the most probable candidate sentence based on a statistical …

    byu Repository record for Surface Realization Using a Featurized Syntactic Statistical Language Model (opens in a new tab)

  13. Query Segmentation For E-Commerce Sites

    … focus on the web search using Google n-gram frequencies corpus or text retrieval from relational databases. However, this module is also useful in the domain of E-Commerce for product search. In this thesis, we will discuss query segmentation in the context of the E-Commerce area. We …

    iupui Repository record for Query Segmentation For E-Commerce Sites (opens in a new tab)

  14. Photo annotation and retrieval through speech

    … speech-based retrieval we have developed a mixed grammar recognition approach which allows the speech recognition system to construct a single finite-state network combining context-free grammars for recognizing and parsing query carrier phrases and metadata phrases, with an unconstrained …

    mit Repository record for Photo annotation and retrieval through speech (opens in a new tab)

  15. Automatic utterance segmentation in spontaneous speech

    … the incorporation of pause duration and grammar information to the utterance segmentation task. As a result, we obtain an optimal set of parameters for the lower level utterance segmenter, and show that part-of-speech based N-gram language modeling of the spoken words in conjunction with …

    mit Repository record for Automatic utterance segmentation in spontaneous speech (opens in a new tab)

  16. Inferring insulin regimen from clinical notes : using natural language processing techniques to extract data from free text records

    … in outpatient clinical notes. We explore two n-gram models - Logistic Regression and Conditional Random Field and analyze their performance. We also explore models using contextual word representations from the domain specific pretrained language models, character level embeddings and auxillary …

    mit Repository record for Inferring insulin regimen from clinical notes : using natural language processing techniques to extract data from free text records (opens in a new tab)

  17. Efficient Substring Discovery Using Suffix, LCP Array and Algorithm-Architecture Interaction

    … and LCP array as a perfect tool to compute N-grams (substring) in various dimensions. Since past couple of decades there has been significant research on construction of suffix and LCP array. Comparatively the research of properly utilizing this prospective data structures to retrieve the …

    lsu-thes Repository record for Efficient Substring Discovery Using Suffix, LCP Array and Algorithm-Architecture Interaction (opens in a new tab)

  18. There Is Always an Option

    … I develop interpretation methods based on $n$-gram statistics of action sequences and mean+variance spatial mappings of agent states. These analyses show that the clusters match clear behaviours and movement patterns when the agent follows an option. Overall, the tool enables the detection of …

    queens Repository record for There Is Always an Option (opens in a new tab)

  19. Evaluation of Automatic Text Summarization Using Synthetic Facts

    … correlation with human judgment with existing N-gram overlap-based metrics such as ROUGE and BLEU and a BERT-based evaluation metric, BERTScore. Our system's experimental evaluation of PEGASUS, BART, and T5 outperforms the current evaluation metrics in measuring factual consistency with a …

    calpoly Repository record for Evaluation of Automatic Text Summarization Using Synthetic Facts (opens in a new tab)

Page 1 of 3