Back to search

York University

Connecting Text and Charts Using Large Vision-Language Models

Abstract

dc:description.abstract

Data visualizations are essential for presenting complex dataset, but the disconnect between charts and accompanying textual descriptions often leads to misinterpretation and increased cognitive effort—especially for users with limited data literacy. While prior methods attempt to bridge this gap, many depend on manual annotations or fixed chart structures, limiting scalability across diverse documents. In this thesis, we propose two large vision-language model (LVLM)-based frameworks—a single-agent baseline and a multi-agent architecture—for automatically linking textual descriptions with their corresponding chart data. Both frameworks extract structured data from chart images and use lexical, syntactic, and arithmetic reasoning to perform sentence-to-data alignment. We evaluate the performance of these frameworks on a curated dataset of Pew Research charts. Finally, we develop a browser extension that integrates this approach into Pew Research articles, enabling interactive text–chart linking for enhanced reading experiences.

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Chowdhury, Nafis Tahmid
Advisor dc:contributor.advisor
  • Prince, Enamul Hoque

Subjects

dc:subject × 1

Rights

dc:rights
Statement dc:rights
  • Author owns copyright, except where explicitly noted. Please contact the author directly with licensing requests.
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier.uri
https://hdl.handle.net/10315/43845

Chain of custody

source
Harvested from
York University
Base URL
yorkspace.library.yorku.ca/oai/request
Last updated
2026-08-21
Source record
OAI-PMH GetRecord
related terms
citation

Chowdhury, Nafis Tahmid. Connecting Text and Charts Using Large Vision-Language Models. 2026. https://hdl.handle.net/10315/43845