Global ETD Search

Search theses and dissertations gathered from participating repositories worldwide. Every result links back to the library that holds it. No account is needed.

Results

Showing 1 to 20 of 34 for “"finetuning"”.

  1. Automated Finetuning via Sparse Autoencoders

    … performance in small models via instruction finetuning. Specifically, we present UnderstandTune, an autonomous method for assembling high-quality instruction finetuning datasets with minimal human intervention, requiring only concise task descriptions rather than evaluation dataset …

    mit Repository record for Automated Finetuning via Sparse Autoencoders (opens in a new tab)

  2. Comparing Parameter Efficient Finetuning Techniques (PEFT) using Datamodels

    … and costly, especially due to the extensive finetuning required. Parameter-efficient finetuning techniques (PEFT) have been proposed to address this issue by significantly reducing the number of trainable parameters, achieving comparable results to full-parameter finetuning. Despite …

    mit Repository record for Comparing Parameter Efficient Finetuning Techniques (PEFT) using Datamodels (opens in a new tab)

  3. Learning to Interpret Language Model Diffs

    Finetuning-induced changes to a model’s weights (a “model diff”) are semantically meaningful but often difficult to interpret. This makes us wonder: can we describe the content of an unknown model diff using natural language? We introduce diff interpretation training, a method that teaches a model …

    mit Repository record for Learning to Interpret Language Model Diffs (opens in a new tab)

  4. Steerable Alignment with Conditional Multiobjective Preference Optimization

    … (RLHF) have provided useful paradigms for finetuning LLMs to produce outputs that are more consistent with human preferences. These approaches, however, assume that preferences are formed by a single, underlying reward model, which is likely insufficient for representing an individual’s …

    mit Repository record for Steerable Alignment with Conditional Multiobjective Preference Optimization (opens in a new tab)

  5. Neural Network Pruning for ECG Arrhythmia Classification

    … employs a pruning phase interleaved with a finetuning phase. It is shown that when performing the scale-factor pruning algorithm on ECG, finetuning time can be expedited by 1.4 times over the traditional approach with only 10% of expensive floating-point operations retained, while …

    calpoly Repository record for Neural Network Pruning for ECG Arrhythmia Classification (opens in a new tab)

  6. Relation extraction: exploring syntax parsing and constructing it as attention-like structure

    … tree structure to matrixes and apply on the finetuning process of SyntaxBERT [2] and syntax-aware -local-attention attention BERT (SLA) [3] to strengthen their ability to learning entity relations. They would be finetuned on TACRED [4] dataset. They are compared with the finetuning of …

    uiuc Repository record for Relation extraction: exploring syntax parsing and constructing it as attention-like structure (opens in a new tab)

  7. Data-Efficient Learning in Image Synthesis and Instance Segmentation

    … high-quality and diverse image generation from finetuning to only 5-100 images. Our method factors a pretrained model into a small but highly expressive weight space for finetuning, which discourages overfitting in a small training set. We validate our method in a challenging few-shot setting of …

    vt Repository record for Data-Efficient Learning in Image Synthesis and Instance Segmentation (opens in a new tab)

  8. Inference Time Search for Protein Structure Prediction

    … search by adding architectural components and a finetuning procedure to state-of-the-art structure prediction models that give rise to a discrete latent space. We implement algorithms for searching and sampling in this discrete latent space and conduct experiments on a small model, demonstrating …

    mit Repository record for Inference Time Search for Protein Structure Prediction (opens in a new tab)

  9. Understanding the Robustness of Vision Models and Humans to Occlusion-Based Corruptions

    … can be mitigated through two approaches: finetuning using occluded images and inpainting occluded pixels before classification. We discover that finetuning leads to a considerable increase in accuracy, but we suspect that finetuned models are relying on a different set of features. …

    mit Repository record for Understanding the Robustness of Vision Models and Humans to Occlusion-Based Corruptions (opens in a new tab)

  10. Adapting Transformers for Structured Data Domains

    … present RAFT-S3, a framework for reasoning-aware finetuning of small language models (SLMs) on the text-to-SQL task. RAFT-S3 collects synthetic text-to-SQL data with diverse schemas using large language models (LLMs), along with intermediate reasoning traces which are incorporated into the …

    vt Repository record for Adapting Transformers for Structured Data Domains (opens in a new tab)

  11. ScPEFT : a parameter-efficient fine-tuning framework for enhancing single-cell large language models in out-of-context

    … outperformed zero-shot models and traditional finetuning in disease-specific, cross-species, and under-characterized cell population tasks. Its attention-mechanism analysis identified COVID-related genes associated with specific cell states and uncovered unique blood cell subpopulations, …

    missouri Repository record for ScPEFT : a parameter-efficient fine-tuning framework for enhancing single-cell large language models in out-of-context (opens in a new tab)

  12. Evaluating Data Augmentation with Attention Masks for Context Aware Transformations

    … augmentations which can be used for further finetuning. Our comprehensive analysis points to limited success of utilizing this context-aware augmentation method. By shedding light on its strengths and limitations, we offer insights that can guide the selection of optimal augmentation …

    mit Repository record for Evaluating Data Augmentation with Attention Masks for Context Aware Transformations (opens in a new tab)

  13. Steering Robots with Inference-Time Interactions

    … behavior. While collecting additional data for finetuning can address such issues, doing so for each downstream use case is inefficient at deployment. My research proposes an alternative: keeping pretrained policies frozen as a fixed skill repertoire while allowing user interactions to guide …

    mit Repository record for Steering Robots with Inference-Time Interactions (opens in a new tab)

  14. Enhanced Potts Models for Improved Computational Protein Design

    … loss function to successfully perform finetuning to improve TERMinator’s performance on orthogonal energetic benchmarks. Finally, I detail an observed disconnect between accuracy on energetic benchmarks and native sequence recovery, illustrating the deficiency of only using native …

    mit Repository record for Enhanced Potts Models for Improved Computational Protein Design (opens in a new tab)

  15. Mixed-Variable Bayesian Optimization using Prior-Data Fitted Networks

    … optimization. Additionally, we explore how finetuning PFNs on targeted function priors can enhance performance when prior knowledge about the objective is available. Our contributions include empirical evaluations of mixed-BO techniques, insights into PFN training, and a suite of …

    mit Repository record for Mixed-Variable Bayesian Optimization using Prior-Data Fitted Networks (opens in a new tab)

  16. ResearchBuddy AI: LLM-Powered Assistant

    … project has three essential features, including finetuning of large language models for structured research paper summarization, prompt engineering for flexible summarization of variable formats of GitHub READMEs, and Retrieval-Augmented Generation (RAG) for context-aware question answering over …

    texas-state Repository record for ResearchBuddy AI: LLM-Powered Assistant (opens in a new tab)

  17. Embodied multimodal referring expressions generation

    … to generate more data for model training; 5) Finetuning LLaMA 2-chat-13B for generating contextually-correct and situationally-fluent multimodal referring expressions; 6) Integrating the fine-tuned model into the IVA to evaluate the success of the generative model-enabled IVA in communication …

    colostate Repository record for Embodied multimodal referring expressions generation (opens in a new tab)

  18. Building a Language Conditioned System for 6-DoF Tabletop Manipulation

    … to either train an end-to-end model or for finetuning. Further, the recent advancements in large models such as Segment Anything and GPT-4 made it possible to construct a modular system, that incorporates vast common sense knowledge, as opposed to traditional approaches. These large models …

    mit Repository record for Building a Language Conditioned System for 6-DoF Tabletop Manipulation (opens in a new tab)

  19. Emergent Capabilities of Generative Models: “Software 3.0” and Beyond

    … programmers are tasked with orchestrating and finetuning the interactions between large-scale foundation models to carry out higher-order tasks.

    mit Repository record for Emergent Capabilities of Generative Models: “Software 3.0” and Beyond (opens in a new tab)

Page 1 of 2