Abstract
dc:descriptionRetrieval-Augmented Generation (RAG) is a technique to augment language models with external knowledge of corpus. Despite the rapid evolution of large language models, RAG is still a promising method for solving the difficulty of updating information and unreliable memorization of large language models as many research endeavors and commercial services leveraged retrieval-augmented generation to improve reliability. However, RAG has its drawbacks including high latency and intensive computational resource utilization. The inefficiency resides in two aspects: the long input due to retrieved documents and slow autoregressive generation. To address these two issues, we propose Efficient Title Reranker, a fast reranker to select important documents for input, and Cascade Speculative Drafting which improves upon speculative decoding to increase the generation efficiency of large language models. The Efficient Title Reranker achieves state-of-the-art in retrieval accuracy while being more efficient than the baseline on the KILT knowledge benchmark. On the other hand, Cascade Speculative Drafting outperforms Speculative Decoding in generation speed on both GSM8k and MMLU without additional training.
Degree
thesis:*- Name thesis:degree_name
- M.S.
- Level thesis:degree_level
- Thesis
- Discipline thesis:degree_discipline
- Computer Science
- Grantor
- University of Illinois at Urbana-Champaign
- Year dc:date
- 2024
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Chen, Ziyi
- Contributors dc:contributor
-
- Chang, Kevin Chen-Chuan
Subjects
dc:subject × 2Rights
dc:rights- Statement dc:rights
-
- Copyright 2024 Ziyi Chen
- Language dc:language
- en, eng
Identifiers
dc:identifier.*- Handle dc:identifier
- https://hdl.handle.net/2142/124428