{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/132791"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/132791","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Toward AI-augmented data analysis: challenges and opportunities","abstract":"Real-world data analysis remains challenging for many users, especially domain experts, because it involves heterogeneous data formats, complex multi-step processing pipelines, and deep technical expertise. Although recent advances in large language models have motivated systems for natural language to SQL translation, semantic query execution, and agentic data retrieval, these systems remain limited to simple analytical tasks over standard data modalities. This thesis systematically investigates the limitations of AI-assisted data analysis along two critical dimensions: (1) data complexity and (2) analytic complexity. Specifically, it evaluates how well current AI systems handle data in complex, irregular forms and how reliably they can execute analytical workflows that move beyond straightforward SQL query translation. For data complexity, we introduce Chart2CSV, a benchmark of 812 real-world scientific charts paired with expert-validated ground-truth tables, and show that state-of-the-art vision language models misinterpret nearly half of the data points. For analytic complexity, we introduce REPRO-Bench, a benchmark of 112 social science reproducibility tasks, and demonstrate that existing AI agents achieve at most 21.4 percent accuracy. Even with our improved system, REPRO-Agent, performance remains far from adequate for practical use. Together, these results show that existing AI systems lack the perceptual, reasoning, and multi-step planning capabilities necessary for reliable real-world data analysis, highlighting substantial open challenges for future research.","abstract_html":"Real-world data analysis remains challenging for many users, especially domain experts, because it involves heterogeneous data formats, complex multi-step processing pipelines, and deep technical expertise. Although recent advances in large language models have motivated systems for natural language to SQL translation, semantic query execution, and agentic data retrieval, these systems remain limited to simple analytical tasks over standard data modalities. This thesis systematically investigates the limitations of AI-assisted data analysis along two critical dimensions: (1) data complexity and (2) analytic complexity. Specifically, it evaluates how well current AI systems handle data in complex, irregular forms and how reliably they can execute analytical workflows that move beyond straightforward SQL query translation. For data complexity, we introduce Chart2CSV, a benchmark of 812 real-world scientific charts paired with expert-validated ground-truth tables, and show that state-of-the-art vision language models misinterpret nearly half of the data points. For analytic complexity, we introduce REPRO-Bench, a benchmark of 112 social science reproducibility tasks, and demonstrate that existing AI agents achieve at most 21.4 percent accuracy. Even with our improved system, REPRO-Agent, performance remains far from adequate for practical use. Together, these results show that existing AI systems lack the perceptual, reasoning, and multi-step planning capabilities necessary for reliable real-world data analysis, highlighting substantial open challenges for future research.","abstract_has_math":false,"creators":["Hu, Chuxuan"],"institution":"University of Illinois Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Kang, Daniel"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2025,"date_issued":"2025-12","date_published":"2025-12","updated_at":"2026-07-22T22:25:07Z","subjects":["Databases","Artificial Intelligence"],"languages":["en"],"rights":["Copyright 2025 Chuxuan Hu"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/2142/132791","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Kang, Daniel"]},{"key":"dc:creator","label":"Author","values":["Hu, Chuxuan"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2025-12","2025-12-03"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Databases","Artificial Intelligence"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2025 Chuxuan Hu"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://hdl.handle.net/2142/132791"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Real-world data analysis remains challenging for many users, especially domain experts, because it involves heterogeneous data formats, complex multi-step processing pipelines, and deep technical expertise. Although recent advances in large language models have motivated systems for natural language to SQL translation, semantic query execution, and agentic data retrieval, these systems remain limited to simple analytical tasks over standard data modalities. This thesis systematically investigates the limitations of AI-assisted data analysis along two critical dimensions: (1) data complexity and (2) analytic complexity. Specifically, it evaluates how well current AI systems handle data in complex, irregular forms and how reliably they can execute analytical workflows that move beyond straightforward SQL query translation. For data complexity, we introduce Chart2CSV, a benchmark of 812 real-world scientific charts paired with expert-validated ground-truth tables, and show that state-of-the-art vision language models misinterpret nearly half of the data points. For analytic complexity, we introduce REPRO-Bench, a benchmark of 112 social science reproducibility tasks, and demonstrate that existing AI agents achieve at most 21.4 percent accuracy. Even with our improved system, REPRO-Agent, performance remains far from adequate for practical use. Together, these results show that existing AI systems lack the perceptual, reasoning, and multi-step planning capabilities necessary for reliable real-world data analysis, highlighting substantial open challenges for future research.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-12-01","The student, Chuxuan Hu, accepted the attached license on 2025-12-02 at 16:07.","The student, Chuxuan Hu, submitted this Thesis for approval on 2025-12-02 at 17:12.","This Thesis was approved for publication on 2025-12-03 at 12:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23028 on 2026-02-19 at 20:09:51"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Toward AI-augmented data analysis: challenges and opportunities"]}]}],"canonical_facts":{"dc:contributor":["Kang, Daniel"],"dc:creator":["Hu, Chuxuan"],"dc:date":["2025-12","2025-12-03"],"dc:description":["Real-world data analysis remains challenging for many users, especially domain experts, because it involves heterogeneous data formats, complex multi-step processing pipelines, and deep technical expertise. Although recent advances in large language models have motivated systems for natural language to SQL translation, semantic query execution, and agentic data retrieval, these systems remain limited to simple analytical tasks over standard data modalities. This thesis systematically investigates the limitations of AI-assisted data analysis along two critical dimensions: (1) data complexity and (2) analytic complexity. Specifically, it evaluates how well current AI systems handle data in complex, irregular forms and how reliably they can execute analytical workflows that move beyond straightforward SQL query translation. For data complexity, we introduce Chart2CSV, a benchmark of 812 real-world scientific charts paired with expert-validated ground-truth tables, and show that state-of-the-art vision language models misinterpret nearly half of the data points. For analytic complexity, we introduce REPRO-Bench, a benchmark of 112 social science reproducibility tasks, and demonstrate that existing AI agents achieve at most 21.4 percent accuracy. Even with our improved system, REPRO-Agent, performance remains far from adequate for practical use. Together, these results show that existing AI systems lack the perceptual, reasoning, and multi-step planning capabilities necessary for reliable real-world data analysis, highlighting substantial open challenges for future research.","Submission published under a 24 month embargo labeled 'Closed Access', the embargo will last until 2027-12-01","The student, Chuxuan Hu, accepted the attached license on 2025-12-02 at 16:07.","The student, Chuxuan Hu, submitted this Thesis for approval on 2025-12-02 at 17:12.","This Thesis was approved for publication on 2025-12-03 at 12:09.","DSpace SAF Submission Ingestion Package generated from Vireo submission #23028 on 2026-02-19 at 20:09:51"],"dc:format":["application/pdf"],"dc:identifier":["https://hdl.handle.net/2142/132791"],"dc:language":["en"],"dc:rights":["Copyright 2025 Chuxuan Hu"],"dc:subject":["Databases","Artificial Intelligence"],"dc:title":["Toward AI-augmented data analysis: challenges and opportunities"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:07Z"}