{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/113023"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/113023","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Large scale urban patterns in NYC: traffic prediction and analysis via clustering and low rank approximation","abstract":"Traffic management is one of the persistent challenges of the modern industrialized world.It simultaneously reflects both a critical infrastructural necessity and a problem involving a wide range of scales and interactions. Like many sectors and spaces, transportation systems can benefit from novel solutions that utilize Machine Learning approaches to extract useful features and patterns from historical Data. Our interest here is to provide a computational framework to perform traffic prediction and analysis via clustering and low rank approximation. Our work consists of several dependent, large-scale optimization problems in order to estimate the number of taxi passengers who travel from a certain origin node to a certain destination node at any given time of day. A unique penalty term is constructed to guarantee geographic continuity among the passenger count predictions and impose dependency within the set of optimization problems. Furthermore, a framework of hypothesis testing is designed to test and validate the algorithm assumptions and justify the importance of the penalty term. We apply our algorithm and procedures to the large scale NYC Taxi GPS data set for the year of 2018. In Chapter 1 we outline our motivation behind this research and provide a summary of the existing literature in urban traffic and transportation systems space.In Chapter 2 we state and define the question of interest and form our configured Non-Negative Matrix Factorization (NMF) algorithm with various considerations of regularization and geographic continuity penalty.In Chapter 3 we discuss the geographic continuity regularization in more detail, which is a core contribution in this research. We define and establish the intuition behind the penalty term, as well as its quantitative construct. Moreover, we shape a discussion around the choice of distance measure and its implication.In Chapter 4 we go over the data elements used in this research, provide background on the taxi GPS data that is made publicly available by the city of New York, and most importantly, discuss the formation of origin/destination nodes through K-means clustering algorithm. Moreover we demonstrate the core methodology that we have designed and implemented to perform traffic prediction and pattern recognition via clustering and low-rank estimation.In Chapter 5 we present the results of our work and begin to evaluate the Machine Learning framework associated with it. This includes model evaluation based on hyper-parameter configuration, error analysis and convergence, actual vs. predicted number of passengers,and training v.s validation loss.In Chapter 6 we define and establish a unique set of hypothesis tests to test and validate the assumption of geographic continuity in the original data points, and confirm the need for the incorporation of the geographic continuity penalty in the algorithm.Finally, Chapter 7 provides our concluding remarks and summary.","abstract_html":"Traffic management is one of the persistent challenges of the modern industrialized world.It simultaneously reflects both a critical infrastructural necessity and a problem involving a wide range of scales and interactions. Like many sectors and spaces, transportation systems can benefit from novel solutions that utilize Machine Learning approaches to extract useful features and patterns from historical Data. Our interest here is to provide a computational framework to perform traffic prediction and analysis via clustering and low rank approximation. Our work consists of several dependent, large-scale optimization problems in order to estimate the number of taxi passengers who travel from a certain origin node to a certain destination node at any given time of day. A unique penalty term is constructed to guarantee geographic continuity among the passenger count predictions and impose dependency within the set of optimization problems. Furthermore, a framework of hypothesis testing is designed to test and validate the algorithm assumptions and justify the importance of the penalty term. We apply our algorithm and procedures to the large scale NYC Taxi GPS data set for the year of 2018. In Chapter 1 we outline our motivation behind this research and provide a summary of the existing literature in urban traffic and transportation systems space.In Chapter 2 we state and define the question of interest and form our configured Non-Negative Matrix Factorization (NMF) algorithm with various considerations of regularization and geographic continuity penalty.In Chapter 3 we discuss the geographic continuity regularization in more detail, which is a core contribution in this research. We define and establish the intuition behind the penalty term, as well as its quantitative construct. Moreover, we shape a discussion around the choice of distance measure and its implication.In Chapter 4 we go over the data elements used in this research, provide background on the taxi GPS data that is made publicly available by the city of New York, and most importantly, discuss the formation of origin/destination nodes through K-means clustering algorithm. Moreover we demonstrate the core methodology that we have designed and implemented to perform traffic prediction and pattern recognition via clustering and low-rank estimation.In Chapter 5 we present the results of our work and begin to evaluate the Machine Learning framework associated with it. This includes model evaluation based on hyper-parameter configuration, error analysis and convergence, actual vs. predicted number of passengers,and training v.s validation loss.In Chapter 6 we define and establish a unique set of hypothesis tests to test and validate the assumption of geographic continuity in the original data points, and confirm the need for the incorporation of the geographic continuity penalty in the algorithm.Finally, Chapter 7 provides our concluding remarks and summary.","abstract_has_math":false,"creators":["Abolhelm, Marzieh"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Industrial Engineering","degree_department":null,"school":null,"contributors":["Sowers, Richard","Beck, Carolyn","Sun, Ruoyu","DeVille, Lee"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-01-12T21:45:39Z","date_published":"2022-01-12T21:45:39Z","updated_at":"2026-07-22T22:24:52Z","subjects":["Data Analytics, Traffic Prediction, Low Rank Approximation, Pattern Recognition Algorithms, Non-Negative Matrix Factorization, Machine Learning"],"languages":["en"],"rights":["Copyright 2021 Marzieh Abolhelm"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/113023","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Sowers, Richard","Beck, Carolyn","Sun, Ruoyu","DeVille, Lee"]},{"key":"dc:creator","label":"Author","values":["Abolhelm, Marzieh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2022-01-12T21:45:39Z","2021-07-12","2021-08"]},{"key":"dc:type","label":"Dc Type","values":["text","Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Industrial Engineering"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Data Analytics, Traffic Prediction, Low Rank Approximation, Pattern Recognition Algorithms, Non-Negative Matrix Factorization, Machine Learning"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2021 Marzieh Abolhelm"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/113023"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Traffic management is one of the persistent challenges of the modern industrialized world.It simultaneously reflects both a critical infrastructural necessity and a problem involving a wide range of scales and interactions. Like many sectors and spaces, transportation systems can benefit from novel solutions that utilize Machine Learning approaches to extract useful features and patterns from historical Data. Our interest here is to provide a computational framework to perform traffic prediction and analysis via clustering and low rank approximation. Our work consists of several dependent, large-scale optimization problems in order to estimate the number of taxi passengers who travel from a certain origin node to a certain destination node at any given time of day. A unique penalty term is constructed to guarantee geographic continuity among the passenger count predictions and impose dependency within the set of optimization problems. Furthermore, a framework of hypothesis testing is designed to test and validate the algorithm assumptions and justify the importance of the penalty term. We apply our algorithm and procedures to the large scale NYC Taxi GPS data set for the year of 2018. In Chapter 1 we outline our motivation behind this research and provide a summary of the existing literature in urban traffic and transportation systems space.In Chapter 2 we state and define the question of interest and form our configured Non-Negative Matrix Factorization (NMF) algorithm with various considerations of regularization and geographic continuity penalty.In Chapter 3 we discuss the geographic continuity regularization in more detail, which is a core contribution in this research. We define and establish the intuition behind the penalty term, as well as its quantitative construct. Moreover, we shape a discussion around the choice of distance measure and its implication.In Chapter 4 we go over the data elements used in this research, provide background on the taxi GPS data that is made publicly available by the city of New York, and most importantly, discuss the formation of origin/destination nodes through K-means clustering algorithm. Moreover we demonstrate the core methodology that we have designed and implemented to perform traffic prediction and pattern recognition via clustering and low-rank estimation.In Chapter 5 we present the results of our work and begin to evaluate the Machine Learning framework associated with it. This includes model evaluation based on hyper-parameter configuration, error analysis and convergence, actual vs. predicted number of passengers,and training v.s validation loss.In Chapter 6 we define and establish a unique set of hypothesis tests to test and validate the assumption of geographic continuity in the original data points, and confirm the need for the incorporation of the geographic continuity penalty in the algorithm.Finally, Chapter 7 provides our concluding remarks and summary.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Marzieh Abolhelm, accepted the attached license on 2021-07-12 at 14:09.","The student, Marzieh Abolhelm, submitted this Dissertation for approval on 2021-07-12 at 14:29.","This Dissertation was approved for publication on 2021-07-12 at 17:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16870 on 2022-01-12 at 12:45:00","Made available in DSpace on 2022-01-12T21:45:39Z (GMT). No. of bitstreams: 2 ABOLHELM-DISSERTATION-2021.pdf: 5399897 bytes, checksum: 4fcd957017dec8a038b8bc855523d1bb (MD5) LICENSE.txt: 4213 bytes, checksum: 9e5dda5b9cf081ca000e957d59eb3b9b (MD5) Previous issue date: 2021-07-12"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Large scale urban patterns in NYC: traffic prediction and analysis via clustering and low rank approximation"]}]}],"canonical_facts":{"dc:contributor":["Sowers, Richard","Beck, Carolyn","Sun, Ruoyu","DeVille, Lee"],"dc:creator":["Abolhelm, Marzieh"],"dc:date":["2022-01-12T21:45:39Z","2021-07-12","2021-08"],"dc:description":["Traffic management is one of the persistent challenges of the modern industrialized world.It simultaneously reflects both a critical infrastructural necessity and a problem involving a wide range of scales and interactions. Like many sectors and spaces, transportation systems can benefit from novel solutions that utilize Machine Learning approaches to extract useful features and patterns from historical Data. Our interest here is to provide a computational framework to perform traffic prediction and analysis via clustering and low rank approximation. Our work consists of several dependent, large-scale optimization problems in order to estimate the number of taxi passengers who travel from a certain origin node to a certain destination node at any given time of day. A unique penalty term is constructed to guarantee geographic continuity among the passenger count predictions and impose dependency within the set of optimization problems. Furthermore, a framework of hypothesis testing is designed to test and validate the algorithm assumptions and justify the importance of the penalty term. We apply our algorithm and procedures to the large scale NYC Taxi GPS data set for the year of 2018. In Chapter 1 we outline our motivation behind this research and provide a summary of the existing literature in urban traffic and transportation systems space.In Chapter 2 we state and define the question of interest and form our configured Non-Negative Matrix Factorization (NMF) algorithm with various considerations of regularization and geographic continuity penalty.In Chapter 3 we discuss the geographic continuity regularization in more detail, which is a core contribution in this research. We define and establish the intuition behind the penalty term, as well as its quantitative construct. Moreover, we shape a discussion around the choice of distance measure and its implication.In Chapter 4 we go over the data elements used in this research, provide background on the taxi GPS data that is made publicly available by the city of New York, and most importantly, discuss the formation of origin/destination nodes through K-means clustering algorithm. Moreover we demonstrate the core methodology that we have designed and implemented to perform traffic prediction and pattern recognition via clustering and low-rank estimation.In Chapter 5 we present the results of our work and begin to evaluate the Machine Learning framework associated with it. This includes model evaluation based on hyper-parameter configuration, error analysis and convergence, actual vs. predicted number of passengers,and training v.s validation loss.In Chapter 6 we define and establish a unique set of hypothesis tests to test and validate the assumption of geographic continuity in the original data points, and confirm the need for the incorporation of the geographic continuity penalty in the algorithm.Finally, Chapter 7 provides our concluding remarks and summary.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2022-01-12 without embargo terms","The student, Marzieh Abolhelm, accepted the attached license on 2021-07-12 at 14:09.","The student, Marzieh Abolhelm, submitted this Dissertation for approval on 2021-07-12 at 14:29.","This Dissertation was approved for publication on 2021-07-12 at 17:17.","DSpace SAF Submission Ingestion Package generated from Vireo submission #16870 on 2022-01-12 at 12:45:00","Made available in DSpace on 2022-01-12T21:45:39Z (GMT). No. of bitstreams: 2 ABOLHELM-DISSERTATION-2021.pdf: 5399897 bytes, checksum: 4fcd957017dec8a038b8bc855523d1bb (MD5) LICENSE.txt: 4213 bytes, checksum: 9e5dda5b9cf081ca000e957d59eb3b9b (MD5) Previous issue date: 2021-07-12"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/113023"],"dc:language":["en"],"dc:rights":["Copyright 2021 Marzieh Abolhelm"],"dc:subject":["Data Analytics, Traffic Prediction, Low Rank Approximation, Pattern Recognition Algorithms, Non-Negative Matrix Factorization, Machine Learning"],"dc:title":["Large scale urban patterns in NYC: traffic prediction and analysis via clustering and low rank approximation"],"dc:type":["text","Thesis"],"thesis:degree_discipline":["Industrial Engineering"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:24:52Z"}