Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 13 of 13 for “"Text Categorization"”.

  1. Effects of OCR errors on text categorization

    … we report on our experiments on training and categorization of optically recognized documents. In, particular, we present a lexicon-based error correction algorithm to improve the categorization process. This algorithm is based on edit distance techniques and information from highly weighted …

    unlv Repository record for Effects of OCR errors on text categorization (opens in a new tab)

  2. Concept graphs: Applications to biomedical text categorization and concept extraction

    … mines for researchers and practitioners. The text content that makes up these knowledge collections is often unstructured and, thus, extracting relevant or novel information could be nontrivial and costly. In addition, human knowledge and expertise are being transformed into structured digital …

    njit Repository record for Concept graphs: Applications to biomedical text categorization and concept extraction (opens in a new tab)

  3. Thesaurus-aided learning for rule-based categorization of Ocr texts

    … of the rule-based approach to automatic text categorization on OCR collections can be improved by using domain-specific thesauri. A rule-based categorizer was constructed consisting of a C++ program called C-KANT which consults documents and creates a program which can be executed by the …

    unlv Repository record for Thesaurus-aided learning for rule-based categorization of Ocr texts (opens in a new tab)

  4. Towards Accurate and Efficient Classification: A Discriminative and Frequent Pattern-Based Approach

    … impact in a wide range of applications including text categorization, chemical compound classification, software behavior analysis and so on.

    uiuc Repository record for Towards Accurate and Efficient Classification: A Discriminative and Frequent Pattern-Based Approach (opens in a new tab)

  5. Combining Prior Knowledge and Data: Beyond the Bayesian Framework

    We explore this task in three contexts: classification (determining the subject of a newsgroup posting), control (learning to perform tasks such as driving a car up a mountain in simulation), and optimization (optimizing performance of linear algebra operations on different hardware platforms). For …

    uiuc Repository record for Combining Prior Knowledge and Data: Beyond the Bayesian Framework (opens in a new tab)

  6. Decision tree rule-based feature selection for imbalanced data

    … real world applications, e.g., fault diagnosis, text categorization and fraud detection. When dealing with an imbalanced dataset, feature selection becomes an important issue. To address it, this work proposes a feature selection method that is based on a decision tree rule and weighted Gini …

    njit Repository record for Decision tree rule-based feature selection for imbalanced data (opens in a new tab)

  7. Reengineering PhysNet in the uPortal framework

    A Digital Library (DL) is an electronic information storage system focused on meeting the information seeking needs of its constituents. As modern DLs often stay in synchronization with the latest progress of technologies in all fields, interoperability among DLs is often hard to achieve. With the …

    vt Repository record for Reengineering PhysNet in the uPortal framework (opens in a new tab)

  8. Topic mining and categorization in online discussion forums

    Online Forums provide a useful way to engage in discussions about a wide variety of topics, as well as gather custom information for which an exact source may not be available, using a combination of knowledge and human interpretation. Usually forums have categories which cater to a particular …

    uiuc Repository record for Topic mining and categorization in online discussion forums (opens in a new tab)

  9. Computer program categorization with machine learning

    lethbridge

  10. Automatic Detection of Nastiness and Early Signs of Cyberbullying Incidents on Social Media

    … methods to identify extremely aggressive texts automatically. We start by exploiting a wide range of linguistic features to create a machine learning model to detect online abusive content. Then, we build a deep neural architecture to identify offensive content in online short and noisy …

    houston Repository record for Automatic Detection of Nastiness and Early Signs of Cyberbullying Incidents on Social Media (opens in a new tab)

  11. Exploiting knowledge in NLP

    … of topics, entities, concepts, and relations in text. Traditionally, statistical models have been successfully deployed for the aforementioned problems. However, the major trend so far has been: “scaling up by dumbing down”- that is, applying sophisticated statistical algorithms operating on very …

    uiuc Repository record for Exploiting knowledge in NLP (opens in a new tab)

  12. Algoritmos de seleção de características personalizados por classe para categorização de texto

    A categorização de textos é uma importante ferramenta para organização e recuperação de informações em documentos digitais. Uma abordagem comum é representar cada palavra como uma característica. Entretanto, a maior parte das características em um documento textual são irrelevantes para sua …

    brazil-ufpe Repository record for Algoritmos de seleção de características personalizados por classe para categorização de texto (opens in a new tab)