Back to results

Colorado School of Mines. Arthur Lakes Library

Table manners at scale: introducing Lora Lens for efficient analysis & detoxification of language models

Abstract

dc:description.abstract

Transformer-based language models have achieved significant advancements across numerous natural language processing (NLP) tasks. However, as these models grow in scale and complexity, ensuring interpretability and mitigating toxic outputs become increasingly critical challenges. This thesis addresses these issues by first analyzing the role attention heads play in propagating toxicity within models, leveraging a recently proposed interpretability tool known as Attention Lens. By decoding attention head outputs into human-interpretable tokens, we identify specific attention heads contributing disproportionately to toxic content generation and demonstrate that targeted interventions at the head level significantly reduce toxicity without requiring complete model retraining. Despite its effectiveness, Attention Lens faces severe limitations in scalability and computational efficiency. To overcome these limitations, we propose and implement the Lora Lens, an innovative adaptation of Attention Lens employing Low-Rank Adaptation (LoRA) to drastically reduce memory footprint and computational cost. Specifically designed for compatibility with large-scale models such as Llama 3 8B, Lora Lens integrates seamlessly with HuggingFace's transformers library and Microsoft's DeepSpeed framework, enabling efficient distributed training. Our results demonstrate that Lora Lens maintains the interpretative capabilities of Attention Lens while significantly enhancing its efficiency and scalability, allowing practical deployment on models with billions of parameters. Ultimately, this work contributes a practical, scalable interpretability technique, enabling researchers and practitioners to better understand, evaluate, and safely deploy large transformer models.

Degree

thesis:*
Name thesis:degree_name
Master of Science (M.S.)
Level thesis:degree_level
Masters
Discipline thesis:degree_discipline
Applied Mathematics and Statistics
Grantor dc:publisher
Colorado School of Mines. Arthur Lakes Library
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Pettyjohn, Jordan
Advisor dc:contributor.advisor
  • McKenzie, Daniel
Committee members dc:contributor.committeemember
  • Wu Fung, Samy
  • Chard, Kyle
  • Hudson, Nathaniel

Subjects

dc:subject × 5

Rights

dc:rights
Statement dc:rights
  • Copyright of the original work is retained by the author.
Language dc:language.iso
eng, English

Identifiers

dc:identifier.*
Identifier
T 9991
OAI identifier oai:identifier
oai:repository.mines.edu:11124/181185

Chain of custody

source
Harvested from
Colorado School of Mines
Base URL
repository.mines.edu/server/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Pettyjohn, Jordan. Table manners at scale: introducing Lora Lens for efficient analysis & detoxification of language models. Masters thesis, Colorado School of Mines. Arthur Lakes Library, 2025. https://hdl.handle.net/11124/181185