Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 9 of 9 for “"Long Documents"”.

  1. Scholarly Information System for Long Documents and their Elements: Structured Representation and Exploration using Knowledge Graphs and LLMs

    Every year, thousands of students write long and detailed theses and dissertations as part of their graduate requirement. These documents contain valuable knowledge—methods, explanations, figures, and insights—but because they are so long, it is often difficult for researchers to find exactly what …

    vt Repository record for Scholarly Information System for Long Documents and their Elements: Structured Representation and Exploration using Knowledge Graphs and LLMs (opens in a new tab)

  2. Enhancing Layout Understanding via Human-in-the-Loop: A User Study on PDF-to-HTML Conversion for Long Documents

    … document elements, enabling systems that convert documents into searchable and editable formats to enhance accessibility and usability. Nevertheless, the recognition results often contain errors that require manual correction due to small training dataset size, limitations of models, and defects …

    vt Repository record for Enhancing Layout Understanding via Human-in-the-Loop: A User Study on PDF-to-HTML Conversion for Long Documents (opens in a new tab)

  3. Information Extraction of Technical Details From Scholarly Articles

    … progress in information extraction from short documents in the last few years, including social media interaction, news articles, and email excerpts. This research aims to extract technical entities like hardware resources, computing platforms, compute time, programming language, and libraries …

    vt Repository record for Information Extraction of Technical Details From Scholarly Articles (opens in a new tab)

  4. ResearchBuddy AI: LLM-Powered Assistant

    … question answering over domain-specific documents. The research explains model selection through self-hosted open-source models with LoRA (Low Rank Adaptation) and other techniques to achieve cost-effectiveness and control. The research paper summarization capability was developed through …

    texas-state Repository record for ResearchBuddy AI: LLM-Powered Assistant (opens in a new tab)

  5. Neural Document Segmentation Using Weighted Sliding Windows with Transformer Encoders

    The subdivision of documents into semantically coherent segments is a fundamental challenge in Natural Language Processing (NLP), with notable applications in information retrieval and question answering. Effective text segmentation is crucial for enhancing Retrieval-Augmented Generation (RAG) …

    york Repository record for Neural Document Segmentation Using Weighted Sliding Windows with Transformer Encoders (opens in a new tab)

  6. Entity-based long document summarization using LLMs

    Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2026-05-01

    uiuc Repository record for Entity-based long document summarization using LLMs (opens in a new tab)

  7. Improving Abstractive Summarization and Information Consistency Assessment

    … there still remain issues such as how to handle long documents. A key challenge to developing these abstractive summarization systems is assessment. In common with many natural language generation tasks, it is not possible to generate an exhaustive set of reference summarises, so simple …

    cambridge Repository record for Improving Abstractive Summarization and Information Consistency Assessment (opens in a new tab)

  8. Improving the effectiveness of language modeling approaches to information retrieval: bridging the theory-effectiveness gap

    … of general retrieval models has been a long-standing difficult challenge in information retrieval research, yet is also a fundamentally important task, because an improved general retrieval model would benefit every search engine. The language modeling approach to information retrieval …

    uiuc Repository record for Improving the effectiveness of language modeling approaches to information retrieval: bridging the theory-effectiveness gap (opens in a new tab)

  9. Analyzing and Navigating Electronic Theses and Dissertations

    … as well as services to make navigation of long PDF documents easier. In recent years, with advances in the field of machine learning for text data, several techniques have been proposed to support such end-user services. However, limited research has been conducted towards integrating such …

    vt Repository record for Analyzing and Navigating Electronic Theses and Dissertations (opens in a new tab)