Back to results

Virginia Tech

Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding

Abstract

dc:description.abstract

Vision Large Language Models (VLLMs) have demonstrated impressive capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably about complex, dynamic environments and rare but high-risk events (e.g., crashes and near-crashes), whereas existing multimodal benchmarks are dominated by routine scenes, focus on narrow sub-tasks, and often rely on evaluation protocols that are vulnerable to guessing and position bias. To address these gaps, this thesis introduces DVBench, a pioneering benchmark for safety-critical driving video understanding with VLLMs. DVBench is built from naturalistic driving videos in the SHRP~2 study and organized by a three-level hierarchical ability taxonomy aligned with widely adopted ADS scenario frameworks (PEGASUS and NHTSA). It covers 4,000 five-second clips spanning normal driving, crashes, and near-crashes, and approximately 10,000 multiple-choice questions that probe 25 fine-grained perception and reasoning abilities, including environmental conditions, road and lane semantics, hazard assessment, event severity, and fault analysis. To obtain more reliable measurements than standard single-trial evaluation, we further propose GroupEval, which rotates answer positions and requires consistent correctness across permutations to mitigate position bias. Using DVBench and GroupEval, we systematically evaluate 14 state-of-the-art VLLMs (0.5B--72B parameters). No off-the-shelf model exceeds 40% accuracy, and all exhibit substantial gaps between relatively strong low-level perception abilities and much weaker high-level safety reasoning, indicating that current VLLMs are far from deployment-ready as autonomous decision-makers. We also study the effect of injecting textual domain knowledge (e.g., definitions of driving terms) and observe model-dependent but generally modest performance gains. To probe domain adaptation, we conduct a fine-tuning study on Qwen2-VL models using DVBench-style multiple-choice questions. Supervised fine-tuning on 2,800 carefully curated instances yields accuracy gains of up to 10.94 percentage points (up to 43.59% relative) under GroupEval. Notably, a fine-tuned 7B model achieves 36.04% accuracy, surpassing a 72B off-the-shelf baseline while using an order of magnitude fewer parameters. Ability-level analysis shows that fine-tuning particularly improves visually grounded environmental and hazard-related abilities (e.g., atmospheric conditions, lane positioning, risk and severity assessment), while nuanced geometry and maneuver evaluation remain challenging. Overall, this thesis provides: (i) a safety-critical, taxonomy-driven benchmark for driving video understanding; (ii) a robust evaluation strategy for reducing position bias in multiple-choice testing; and (iii) empirical evidence that targeted fine-tuning can substantially narrow---though not close---the gap between general-purpose VLLMs and the stringent requirements of mission-critical driving applications. We release the DVBench benchmark, evaluation toolkit, and fine-tuning scripts and checkpoints at: https://github.com/tong-zeng/DVBench.git, to support reproducible research and future progress in VLLM-based traffic safety understanding.

Degree

thesis:*
Name thesis:degree_name
Master of Science
Level thesis:degree_level
masters
Discipline thesis:degree_discipline
Computer Science & Applications
Department dc:contributor.department
Computer Science and Applications
Grantor dc:publisher
Virginia Tech
Year dc:date.issued
2026

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Zeng, Tong
Chair dc:contributor.committeechair
  • Zhou, Dawei
Committee members dc:contributor.committeemember
  • Guo, Feng
  • North, Christopher L.

Subjects

dc:subject × 7

Rights

dc:rights
Statement dc:rights
  • In Copyright
Language dc:language.iso
en

Identifiers

dc:identifier.*
Dc Identifier Other
vt_gsexam:45572
OAI identifier oai:identifier
oai:vtechworks.lib.vt.edu:10919/140970

Chain of custody

source
Harvested from
Virginia Tech
Base URL
vtechworks.lib.vt.edu/oai/request
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Zeng, Tong. Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding. masters thesis, Virginia Tech, 2026. https://hdl.handle.net/10919/140970