Global ETD Search
Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.
Results
Showing 1 to 20 of 22 for “"document representation"”.
-
Ontology-based document representation for biomedical information retrieval
… of a query term with a term contained in the document representation. Biomedical ontologies, when available, can help disambiguate the information expressed in free-text: they provide unique terms to represent concepts and therefore counterweiglzt the occurrence of synonyms and polysems in …
-
Knowledge-enhanced Document Representation Learning for Legal Judgment Support
… focusing on the enhancement of legal knowledge representation within AI algorithms, aimed at supporting judicial work. This research explores three critical aspects: (1) Classification: The initial step in legal proceedings involves the categorization of submitted documents, a process …
-
Improving document representation by accumulating relevance feedback : the relevance feedback accumulation (RFA) algorithm
Document representation (indexing) techniques are dominated by variants of the term-frequency analysis approach, based on the assumption that the more occurrences a term has throughout a document the more important the term is in that document. Inherent drawbacks associated with this approach …
-
Building an Intelligent Filtering System Using Idea Indexing
… the vector model. The research on improving document representation has been focused on four areas, namely, statistical co-occurrence of related items, forming term phrases, grouping of related words, and representing the content of documents. In this thesis, we propose the idea-indexing …
-
A Study of Graphically Chosen Features for Representation of TREC Topic-Document Sets
Document representation is important for computer-based text processing. Good document representations must include at least the most salient concepts of the document. Documents exist in a multidimensional space that difficult the identification of what concepts to include. A current problem is to …
-
New Weighting Schemes for Document Ranking and Ranked Query Suggestion
… need or the importance of a term to a document. This thesis aims to investigate novel term weighting methods with applications in document representation for text classification, web document ranking, and ranked query suggestion. Firstly, this research proposes a new feature for …
-
QuOTE: Question-Oriented Text Embeddings
… generation (RAG) systems, aimed at improving document representation for accurate and nuanced retrieval. Unlike traditional RAG pipelines, which rely on embed- ding raw text chunks, QuOTE augments chunks with hypothetical questions that the chunk can potentially answer, enriching the …
-
High compression rate text summarization
… thesis focuses on methods for condensing large documents into highly concise summaries, achieving compression rates on par with human writers. While the need for such summaries in the current age of information overload is increasing, the desired compression rate has thus far been beyond the …
-
Software Requirements Classification Using Word Embeddings and Convolutional Neural Networks
… specifically the use of word embeddings for document representation when training a convolutional neural network (CNN). As past research endeavors mainly utilize information retrieval and traditional machine learning techniques, we entertain the potential of deep learning on this particular …
-
Evolutionary learning multi-agent based information retrieval systems
… of matching user interests and the retrieved documents. First, is the fact that users often do not present queries to information retrieval systems in the form that optimally represents the information they want. Secondly, the measure of a document's relevance is highly subjective and variable …
-
Spectral Regression: A Regression Framework for Efficient Regularized Subspace Learning
… real world applications, e.g. face analysis, document representation and content-based image retrieval.
-
A system for document analysis, translation, and automatic hypertext linking
… database is a heterogeneous collection of documents. Documents may become available in different formats (e.g., ASCII, SGML, typesetter languages) and they may have to be translated to a standard document representation scheme used by the digital library. This work focuses on the design of …
-
Document expansion and language model re-estimation for information retrieval
Document expansion is the process of augmenting the text of a document with text drawn from one or more other documents. The purpose of this expansion is to increase the size of the term sample from which document representations, such as language models, may be estimated. While document expansion …
-
Representation and learning schemes for sentiment analysis.
… extraction and selection, enrichment of the document representation and exploitation of the ordinal structure of rating classes. The techniques were evaluated on four sentiment-rich corpora, using two well-known classifiers: Support Vector Machines and Na¨ıve Bayes. This thesis proposes the …
-
A temporal topic model for social trend prediction
… predictive hidden topics. The model includes document partitioning, topic inference, topic selection, and document representation phases. In fact, a dynamic vocabulary is built to detect emerging topics. The extracted topics are compared over time to select more diverse and novel topics in …
-
Role of semantic indexing for text classification.
The Vector Space Model (VSM) of text representation suffers a number of limitations for text classification. Firstly, the VSM is based on the Bag-Of-Words (BOW) assumption where terms from the indexing vocabulary are treated independently of one another. However, the expressiveness of natural …
-
Information filtering by multiple examples
… their information needs as a set of relevant documents rather than as a set of keywords. Most of the studies on SBME adopt the Positive Unlabeled learning (PU learning) techniques by treating the user's provided examples (denoted as query examples) as positive set and the entire data …
-
Constructing and modeling text-rich information networks: a phrase mining-based approach
… in everyone's daily life in forms of web documents, business reviews, news, social posts, etc. In the mean time, textual data and structured entities often come in intertwined, such as authors/posters, document categories and tags, and document-associated geo locations. With this …
-
Representing semantic relatedness
… question we must address is how to represent documents. The way a document is organised reflects certain explicit and implicit semantic and syntactical coupling relationships which are embedded in its contents. The effective capturing of such content couplings is thereby crucial for a genuine …
-
Arabic Language Processing for Text Classification. Contributions to Arabic Root Extraction Techniques, Building An Arabic Corpus, and to Arabic Text Classification Techniques.
… and implements a variant term frequency inverse document frequency weighting method, and investigates the effect of using different choices of features in document representation on single-label text classification performance (words, stems or roots as well as including to these choices their …
Page 1 of 2