Virginia Tech
Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding
Abstract
dc:description.abstractVision Large Language Models (VLLMs) have demonstrated impressive capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably about complex, dynamic environments and rare but high-risk events (e.g., crashes and near-crashes), whereas existing multimodal benchmarks are dominated by routine scenes, focus on narrow sub-tasks, and often rely on evaluation protocols that are vulnerable to guessing and position bias. To address these gaps, this thesis introduces DVBench, a pioneering benchmark for safety-critical driving video understanding with VLLMs. DVBench is built from naturalistic driving videos in the SHRP~2 study and organized by a three-level hierarchical ability taxonomy aligned with widely adopted ADS scenario frameworks (PEGASUS and NHTSA). It covers 4,000 five-second clips spanning normal driving, crashes, and near-crashes, and approximately 10,000 multiple-choice questions that probe 25 fine-grained perception and reasoning abilities, including environmental conditions, road and lane semantics, hazard assessment, event severity, and fault analysis. To obtain more reliable measurements than standard single-trial evaluation, we further propose GroupEval, which rotates answer positions and requires consistent correctness across permutations to mitigate position bias. Using DVBench and GroupEval, we systematically evaluate 14 state-of-the-art VLLMs (0.5B--72B parameters). No off-the-shelf model exceeds 40% accuracy, and all exhibit substantial gaps between relatively strong low-level perception abilities and much weaker high-level safety reasoning, indicating that current VLLMs are far from deployment-ready as autonomous decision-makers. We also study the effect of injecting textual domain knowledge (e.g., definitions of driving terms) and observe model-dependent but generally modest performance gains. To probe domain adaptation, we conduct a fine-tuning study on Qwen2-VL models using DVBench-style multiple-choice questions. Supervised fine-tuning on 2,800 carefully curated instances yields accuracy gains of up to 10.94 percentage points (up to 43.59% relative) under GroupEval. Notably, a fine-tuned 7B model achieves 36.04% accuracy, surpassing a 72B off-the-shelf baseline while using an order of magnitude fewer parameters. Ability-level analysis shows that fine-tuning particularly improves visually grounded environmental and hazard-related abilities (e.g., atmospheric conditions, lane positioning, risk and severity assessment), while nuanced geometry and maneuver evaluation remain challenging. Overall, this thesis provides: (i) a safety-critical, taxonomy-driven benchmark for driving video understanding; (ii) a robust evaluation strategy for reducing position bias in multiple-choice testing; and (iii) empirical evidence that targeted fine-tuning can substantially narrow---though not close---the gap between general-purpose VLLMs and the stringent requirements of mission-critical driving applications. We release the DVBench benchmark, evaluation toolkit, and fine-tuning scripts and checkpoints at: https://github.com/tong-zeng/DVBench.git, to support reproducible research and future progress in VLLM-based traffic safety understanding.
Degree
thesis:*- Name thesis:degree_name
- Master of Science
- Level thesis:degree_level
- masters
- Discipline thesis:degree_discipline
- Computer Science & Applications
- Department dc:contributor.department
- Computer Science and Applications
- Grantor dc:publisher
- Virginia Tech
- Year dc:date.issued
- 2026
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Zeng, Tong
- Chair dc:contributor.committeechair
-
- Zhou, Dawei
- Committee members dc:contributor.committeemember
-
- Guo, Feng
- North, Christopher L.
Subjects
dc:subject × 7Rights
dc:rights- Statement dc:rights
-
- In Copyright
- Licence dc:rights.uri
- Language dc:language.iso
- en
Identifiers
dc:identifier.*- Dc Identifier Other
- vt_gsexam:45572
- OAI identifier oai:identifier
- oai:vtechworks.lib.vt.edu:10919/140970