{"id":{"repo_id":"vt","oai_identifier":"oai:vtechworks.lib.vt.edu:10919/140970"},"canonical_url":"https://search.dev.ndltd.org/etd/vt/oai:vtechworks.lib.vt.edu:10919/140970","repository":{"repo_id":"vt","name":"Virginia Tech","base_url":"https://vtechworks.lib.vt.edu/oai/request"},"display":{"title":"Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding","abstract":"Vision Large Language Models (VLLMs) have demonstrated impressive capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably about complex, dynamic environments and rare but high-risk events (e.g., crashes and near-crashes), whereas existing multimodal benchmarks are dominated by routine scenes, focus on narrow sub-tasks, and often rely on evaluation protocols that are vulnerable to guessing and position bias. To address these gaps, this thesis introduces DVBench, a pioneering benchmark for safety-critical driving video understanding with VLLMs. DVBench is built from naturalistic driving videos in the SHRP~2 study and organized by a three-level hierarchical ability taxonomy aligned with widely adopted ADS scenario frameworks (PEGASUS and NHTSA). It covers 4,000 five-second clips spanning normal driving, crashes, and near-crashes, and approximately 10,000 multiple-choice questions that probe 25 fine-grained perception and reasoning abilities, including environmental conditions, road and lane semantics, hazard assessment, event severity, and fault analysis. To obtain more reliable measurements than standard single-trial evaluation, we further propose GroupEval, which rotates answer positions and requires consistent correctness across permutations to mitigate position bias. Using DVBench and GroupEval, we systematically evaluate 14 state-of-the-art VLLMs (0.5B--72B parameters). No off-the-shelf model exceeds 40% accuracy, and all exhibit substantial gaps between relatively strong low-level perception abilities and much weaker high-level safety reasoning, indicating that current VLLMs are far from deployment-ready as autonomous decision-makers. We also study the effect of injecting textual domain knowledge (e.g., definitions of driving terms) and observe model-dependent but generally modest performance gains. To probe domain adaptation, we conduct a fine-tuning study on Qwen2-VL models using DVBench-style multiple-choice questions. Supervised fine-tuning on 2,800 carefully curated instances yields accuracy gains of up to 10.94 percentage points (up to 43.59% relative) under GroupEval. Notably, a fine-tuned 7B model achieves 36.04% accuracy, surpassing a 72B off-the-shelf baseline while using an order of magnitude fewer parameters. Ability-level analysis shows that fine-tuning particularly improves visually grounded environmental and hazard-related abilities (e.g., atmospheric conditions, lane positioning, risk and severity assessment), while nuanced geometry and maneuver evaluation remain challenging. Overall, this thesis provides: (i) a safety-critical, taxonomy-driven benchmark for driving video understanding; (ii) a robust evaluation strategy for reducing position bias in multiple-choice testing; and (iii) empirical evidence that targeted fine-tuning can substantially narrow---though not close---the gap between general-purpose VLLMs and the stringent requirements of mission-critical driving applications. We release the DVBench benchmark, evaluation toolkit, and fine-tuning scripts and checkpoints at: https://github.com/tong-zeng/DVBench.git, to support reproducible research and future progress in VLLM-based traffic safety understanding.","abstract_html":"Vision Large Language Models (VLLMs) have demonstrated impressive capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably about complex, dynamic environments and rare but high-risk events (e.g., crashes and near-crashes), whereas existing multimodal benchmarks are dominated by routine scenes, focus on narrow sub-tasks, and often rely on evaluation protocols that are vulnerable to guessing and position bias. To address these gaps, this thesis introduces DVBench, a pioneering benchmark for safety-critical driving video understanding with VLLMs. DVBench is built from naturalistic driving videos in the SHRP~2 study and organized by a three-level hierarchical ability taxonomy aligned with widely adopted ADS scenario frameworks (PEGASUS and NHTSA). It covers 4,000 five-second clips spanning normal driving, crashes, and near-crashes, and approximately 10,000 multiple-choice questions that probe 25 fine-grained perception and reasoning abilities, including environmental conditions, road and lane semantics, hazard assessment, event severity, and fault analysis. To obtain more reliable measurements than standard single-trial evaluation, we further propose GroupEval, which rotates answer positions and requires consistent correctness across permutations to mitigate position bias. Using DVBench and GroupEval, we systematically evaluate 14 state-of-the-art VLLMs (0.5B--72B parameters). No off-the-shelf model exceeds 40% accuracy, and all exhibit substantial gaps between relatively strong low-level perception abilities and much weaker high-level safety reasoning, indicating that current VLLMs are far from deployment-ready as autonomous decision-makers. We also study the effect of injecting textual domain knowledge (e.g., definitions of driving terms) and observe model-dependent but generally modest performance gains. To probe domain adaptation, we conduct a fine-tuning study on Qwen2-VL models using DVBench-style multiple-choice questions. Supervised fine-tuning on 2,800 carefully curated instances yields accuracy gains of up to 10.94 percentage points (up to 43.59% relative) under GroupEval. Notably, a fine-tuned 7B model achieves 36.04% accuracy, surpassing a 72B off-the-shelf baseline while using an order of magnitude fewer parameters. Ability-level analysis shows that fine-tuning particularly improves visually grounded environmental and hazard-related abilities (e.g., atmospheric conditions, lane positioning, risk and severity assessment), while nuanced geometry and maneuver evaluation remain challenging. Overall, this thesis provides: (i) a safety-critical, taxonomy-driven benchmark for driving video understanding; (ii) a robust evaluation strategy for reducing position bias in multiple-choice testing; and (iii) empirical evidence that targeted fine-tuning can substantially narrow---though not close---the gap between general-purpose VLLMs and the stringent requirements of mission-critical driving applications. We release the DVBench benchmark, evaluation toolkit, and fine-tuning scripts and checkpoints at: https://github.com/tong-zeng/DVBench.git, to support reproducible research and future progress in VLLM-based traffic safety understanding.","abstract_has_math":false,"creators":["Zeng, Tong"],"institution":"Virginia Tech","degree_name":"Master of Science","degree_level":"masters","degree_discipline":"Computer Science & Applications","degree_department":"Computer Science and Applications","school":null,"contributors":[],"advisors":[],"committee_chairs":["Zhou, Dawei"],"committee_members":["Guo, Feng","North, Christopher L."],"year":2026,"date_issued":"2026-01-23","date_published":"2026-01-23","updated_at":"2026-07-22T22:20:09Z","subjects":["Driving Video Understanding","Vision Large Language Models","Safety-Critical Events","Autonomous Driving Systems","Multi-Modal Learning","Benchmarking","Domain Adaptation"],"languages":["en"],"rights":["In Copyright"],"rights_urls":["http://rightsstatements.org/vocab/InC/1.0/"],"identifier_entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45572"],"render_values":[{"text":"vt_gsexam:45572","href":null,"code":true}]}]},"links":{"outbound_url":"https://hdl.handle.net/10919/140970","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.committeechair","label":"Committee Chair","values":["Zhou, Dawei"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Guo, Feng","North, Christopher L."]},{"key":"dc:contributor.department","label":"Department","values":["Computer Science and Applications"]},{"key":"dc:creator","label":"Author","values":["Zeng, Tong"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2026-01-24T09:00:10Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2026-01-24T09:00:10Z"]},{"key":"dc:date.issued","label":"Date","values":["2026-01-23"]},{"key":"dc:publisher","label":"Institution","values":["Virginia Tech"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science & Applications"]},{"key":"thesis:degree_level","label":"Degree Level","values":["masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Virginia Polytechnic Institute and State University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Driving Video Understanding","Vision Large Language Models","Safety-Critical Events","Autonomous Driving Systems","Multi-Modal Learning","Benchmarking","Domain Adaptation"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["In Copyright"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/vocab/InC/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.other","label":"Dc Identifier Other","values":["vt_gsexam:45572"]},{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/10919/140970"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Vision Large Language Models (VLLMs) have demonstrated impressive capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably about complex, dynamic environments and rare but high-risk events (e.g., crashes and near-crashes), whereas existing multimodal benchmarks are dominated by routine scenes, focus on narrow sub-tasks, and often rely on evaluation protocols that are vulnerable to guessing and position bias. To address these gaps, this thesis introduces DVBench, a pioneering benchmark for safety-critical driving video understanding with VLLMs. DVBench is built from naturalistic driving videos in the SHRP~2 study and organized by a three-level hierarchical ability taxonomy aligned with widely adopted ADS scenario frameworks (PEGASUS and NHTSA). It covers 4,000 five-second clips spanning normal driving, crashes, and near-crashes, and approximately 10,000 multiple-choice questions that probe 25 fine-grained perception and reasoning abilities, including environmental conditions, road and lane semantics, hazard assessment, event severity, and fault analysis. To obtain more reliable measurements than standard single-trial evaluation, we further propose GroupEval, which rotates answer positions and requires consistent correctness across permutations to mitigate position bias. Using DVBench and GroupEval, we systematically evaluate 14 state-of-the-art VLLMs (0.5B--72B parameters). No off-the-shelf model exceeds 40% accuracy, and all exhibit substantial gaps between relatively strong low-level perception abilities and much weaker high-level safety reasoning, indicating that current VLLMs are far from deployment-ready as autonomous decision-makers. We also study the effect of injecting textual domain knowledge (e.g., definitions of driving terms) and observe model-dependent but generally modest performance gains. To probe domain adaptation, we conduct a fine-tuning study on Qwen2-VL models using DVBench-style multiple-choice questions. Supervised fine-tuning on 2,800 carefully curated instances yields accuracy gains of up to 10.94 percentage points (up to 43.59% relative) under GroupEval. Notably, a fine-tuned 7B model achieves 36.04% accuracy, surpassing a 72B off-the-shelf baseline while using an order of magnitude fewer parameters. Ability-level analysis shows that fine-tuning particularly improves visually grounded environmental and hazard-related abilities (e.g., atmospheric conditions, lane positioning, risk and severity assessment), while nuanced geometry and maneuver evaluation remain challenging. Overall, this thesis provides: (i) a safety-critical, taxonomy-driven benchmark for driving video understanding; (ii) a robust evaluation strategy for reducing position bias in multiple-choice testing; and (iii) empirical evidence that targeted fine-tuning can substantially narrow---though not close---the gap between general-purpose VLLMs and the stringent requirements of mission-critical driving applications. We release the DVBench benchmark, evaluation toolkit, and fine-tuning scripts and checkpoints at: https://github.com/tong-zeng/DVBench.git, to support reproducible research and future progress in VLLM-based traffic safety understanding."]},{"key":"dc:description.abstractgeneral","label":"General Abstract","values":["Every year, road crashes kill or injure millions of people worldwide. Many of these crashes are linked to human mistakes, such as misjudging other vehicles or failing to notice hazards in time. One promise of self-driving and advanced driver-assistance systems is to reduce these errors by using cameras and artificial intelligence (AI) to constantly ``watch'' the road and help vehicles respond more safely. Recently, powerful AI models have been developed that can look at images or videos and describe what they see in words or answer questions about them. These systems are already used for everyday tasks like captioning photos or helping people with visual impairments. However, we still know very little about how well they understand truly dangerous driving situations, such as crashes and near-crashes, which matter most for safety. This thesis studies how well such vision-and-language AI models understand real driving videos, especially during safety-critical moments. I first built a large test collection of short video clips taken from naturalistic driving studies, including normal driving, near-misses, and actual crashes. For each clip, human annotators wrote multiple-choice questions about what is happening, who is at risk, and how dangerous the situation is. I also designed an evaluation method that avoids giving models \"easy points\" just because they prefer certain answer positions (for example, always guessing option A). Using this benchmark, I tested a wide range of state-of-the-art AI models. I found that while they can easily recognize objects like cars and pedestrians, they frequently misjudge risks, causes, and appropriate driver actions. To address this, I explored \"fine-tuning\"—a process of retraining the AI specifically on driving data, similar to sending a general student to a specialized driving school. This specialized training significantly boosted the models' ability to identify environmental hazards and assess risks. Notably, I found that a smaller, fine-tuned model could outperform a massive general-purpose model that was ten times its size. While this narrows the gap between current AI and human reliability, the results show that even the best models are not yet ready to take the wheel in complex emergencies."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Master of Science"]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["ETD"]},{"key":"dc:title","label":"Title","values":["Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding"]}]}],"canonical_facts":{"dc:contributor.committeechair":["Zhou, Dawei"],"dc:contributor.committeemember":["Guo, Feng","North, Christopher L."],"dc:contributor.department":["Computer Science and Applications"],"dc:creator":["Zeng, Tong"],"dc:date.accessioned":["2026-01-24T09:00:10Z"],"dc:date.available":["2026-01-24T09:00:10Z"],"dc:date.issued":["2026-01-23"],"dc:description.abstract":["Vision Large Language Models (VLLMs) have demonstrated impressive capabilities on general-purpose image and video understanding tasks, such as captioning and visual question answering. However, their effectiveness in specialized, safety-critical domains like autonomous driving remains largely unclear. Autonomous Driving Systems (ADS) must reason reliably about complex, dynamic environments and rare but high-risk events (e.g., crashes and near-crashes), whereas existing multimodal benchmarks are dominated by routine scenes, focus on narrow sub-tasks, and often rely on evaluation protocols that are vulnerable to guessing and position bias. To address these gaps, this thesis introduces DVBench, a pioneering benchmark for safety-critical driving video understanding with VLLMs. DVBench is built from naturalistic driving videos in the SHRP~2 study and organized by a three-level hierarchical ability taxonomy aligned with widely adopted ADS scenario frameworks (PEGASUS and NHTSA). It covers 4,000 five-second clips spanning normal driving, crashes, and near-crashes, and approximately 10,000 multiple-choice questions that probe 25 fine-grained perception and reasoning abilities, including environmental conditions, road and lane semantics, hazard assessment, event severity, and fault analysis. To obtain more reliable measurements than standard single-trial evaluation, we further propose GroupEval, which rotates answer positions and requires consistent correctness across permutations to mitigate position bias. Using DVBench and GroupEval, we systematically evaluate 14 state-of-the-art VLLMs (0.5B--72B parameters). No off-the-shelf model exceeds 40% accuracy, and all exhibit substantial gaps between relatively strong low-level perception abilities and much weaker high-level safety reasoning, indicating that current VLLMs are far from deployment-ready as autonomous decision-makers. We also study the effect of injecting textual domain knowledge (e.g., definitions of driving terms) and observe model-dependent but generally modest performance gains. To probe domain adaptation, we conduct a fine-tuning study on Qwen2-VL models using DVBench-style multiple-choice questions. Supervised fine-tuning on 2,800 carefully curated instances yields accuracy gains of up to 10.94 percentage points (up to 43.59% relative) under GroupEval. Notably, a fine-tuned 7B model achieves 36.04% accuracy, surpassing a 72B off-the-shelf baseline while using an order of magnitude fewer parameters. Ability-level analysis shows that fine-tuning particularly improves visually grounded environmental and hazard-related abilities (e.g., atmospheric conditions, lane positioning, risk and severity assessment), while nuanced geometry and maneuver evaluation remain challenging. Overall, this thesis provides: (i) a safety-critical, taxonomy-driven benchmark for driving video understanding; (ii) a robust evaluation strategy for reducing position bias in multiple-choice testing; and (iii) empirical evidence that targeted fine-tuning can substantially narrow---though not close---the gap between general-purpose VLLMs and the stringent requirements of mission-critical driving applications. We release the DVBench benchmark, evaluation toolkit, and fine-tuning scripts and checkpoints at: https://github.com/tong-zeng/DVBench.git, to support reproducible research and future progress in VLLM-based traffic safety understanding."],"dc:description.abstractgeneral":["Every year, road crashes kill or injure millions of people worldwide. Many of these crashes are linked to human mistakes, such as misjudging other vehicles or failing to notice hazards in time. One promise of self-driving and advanced driver-assistance systems is to reduce these errors by using cameras and artificial intelligence (AI) to constantly ``watch'' the road and help vehicles respond more safely. Recently, powerful AI models have been developed that can look at images or videos and describe what they see in words or answer questions about them. These systems are already used for everyday tasks like captioning photos or helping people with visual impairments. However, we still know very little about how well they understand truly dangerous driving situations, such as crashes and near-crashes, which matter most for safety. This thesis studies how well such vision-and-language AI models understand real driving videos, especially during safety-critical moments. I first built a large test collection of short video clips taken from naturalistic driving studies, including normal driving, near-misses, and actual crashes. For each clip, human annotators wrote multiple-choice questions about what is happening, who is at risk, and how dangerous the situation is. I also designed an evaluation method that avoids giving models \"easy points\" just because they prefer certain answer positions (for example, always guessing option A). Using this benchmark, I tested a wide range of state-of-the-art AI models. I found that while they can easily recognize objects like cars and pedestrians, they frequently misjudge risks, causes, and appropriate driver actions. To address this, I explored \"fine-tuning\"—a process of retraining the AI specifically on driving data, similar to sending a general student to a specialized driving school. This specialized training significantly boosted the models' ability to identify environmental hazards and assess risks. Notably, I found that a smaller, fine-tuned model could outperform a massive general-purpose model that was ten times its size. While this narrows the gap between current AI and human reliability, the results show that even the best models are not yet ready to take the wheel in complex emergencies."],"dc:description.degree":["Master of Science"],"dc:format.medium":["ETD"],"dc:identifier.other":["vt_gsexam:45572"],"dc:identifier.uri":["https://hdl.handle.net/10919/140970"],"dc:language.iso":["en"],"dc:publisher":["Virginia Tech"],"dc:rights":["In Copyright"],"dc:rights.uri":["http://rightsstatements.org/vocab/InC/1.0/"],"dc:subject":["Driving Video Understanding","Vision Large Language Models","Safety-Critical Events","Autonomous Driving Systems","Multi-Modal Learning","Benchmarking","Domain Adaptation"],"dc:title":["Are Vision Large Language Models Road-Ready? Benchmarking and Adapting VLLMs for Safety-Critical Driving Video Understanding"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science & Applications"],"thesis:degree_level":["masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Virginia Polytechnic Institute and State University"]},"updated_at":"2026-07-22T22:20:09Z"}