Back to results

Virginia Tech

Model Control through Lightweight Activation Steering for Vision Language Models

Abstract

dc:description.abstract

This work introduces SteerVLM, a lightweight steering module designed to guide Vision- Language Models (VLMs) towards outputs that better adhere to desired instructions. Our approach learns from the latent embeddings of paired prompts encoding target and converse behaviors to dynamically adjust activations connecting the language modality with image context. This provides fine-grained, inference-time control over complex output semantics without modifying model weights while preserving performance on off-target tasks. Our steering module requires learning parameters equal to 0.14% of the original VLM's size. Additionally, our steering module gains model control via dimension-wise activation mod- ulation and adaptive layer-wise steering without requiring pre-extracted static vectors or manual tuning of intervention points. Furthermore, we introduce VNIA (Visual Narrative Intent Alignment), a multimodal dataset specifically created to facilitate the development and evaluation of VLM steering techniques. Our method outperforms existing intervention techniques on steering and hallucination mitigation benchmarks for VLMs and proposes a robust solution for multimodal model control through activation engineering.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Level thesis:degree_level
masters
Discipline thesis:degree_discipline
Computer Science & Applications
Department dc:contributor.department
Computer Science and#38; Applications
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2025

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Sivakumar, Anushka
Chair dc:contributor.committeechair
  • Thomas, Christopher Lee
Committee members dc:contributor.committeemember
  • Wang, Xuan
  • Yanardag Delul, Pinar

Subjects

dc:subject × 6

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:44270
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/135566

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-24
Source record
OAI-PMH GetRecord
citation

Sivakumar, Anushka. Model Control through Lightweight Activation Steering for Vision Language Models. masters thesis, Virginia Tech, 2025. https://hdl.handle.net/10919/135566