{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/95338"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/95338","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Modeling the winning seed distribution of the NCAA basketball tournament","abstract":"The National Collegiate Athletic Association's (NCAA) men's division I college basketball tournament is an annual competition that draws widespread attention in the United States. Estimating the outcome of each game is a popular activity undertaken by numerous websites, fans, and more recently, academic researchers. There has been a surge of interest in proposing mathematical methods to model the tournament's results and pick the winners of future games. This thesis analyzes the results of the NCAA basketball tournament since 1985 and proposes several models to capture the winning seed distribution in each round. The Exponential Model estimates the winning probability of each team by modeling the time between a team's successive winnings in a round as an exponential random variable. The Exponential Model estimates a zero probability for events that have not occurred in the training data set. The Markov Model solves this limitation by defining a Markov chain that incorporates each team's winnings in prior rounds to estimate its winning probability. Results of these two models are validated using a chi-squared goodness of fit test. The Power Model, which is an intelligent tool for generating brackets of winners, quantifies the relative strength of each match-up in a round as a power function of the teams' seed numbers, with the exponent estimated using the historical results. The main problem of the Power Model is the data complications that are generally caused by the small size of the training data set, especially in later rounds. The Position and Upset Models solve this problem by representing the tournament's games as a binary sequence and estimating the outcome of each game based on the teams' performance in the similar game. While generating a bracket in a forward direction from the first to the last round propagates the incorrect picks through the tournament, correctly picking the winners in later rounds automatically fills the bracket for several games in earlier rounds. This motivates developing bidirectional models that pick the winners based on a combination of models in forward and backward directions. The Power, Position, Upset, and bidirectional models are assessed based on the aggregate performance of millions of brackets for the five most recent tournaments (2012-2016). The proposed models allow one to estimate the likelihoods of different seed combinations by applying the estimated winning seed distributions, which accurately summarize the seeds' aggregate performance and provide a deeper understanding of the uncertainty in the games' outcomes.","abstract_html":"The National Collegiate Athletic Association&#x27;s (NCAA) men&#x27;s division I college basketball tournament is an annual competition that draws widespread attention in the United States. Estimating the outcome of each game is a popular activity undertaken by numerous websites, fans, and more recently, academic researchers. There has been a surge of interest in proposing mathematical methods to model the tournament&#x27;s results and pick the winners of future games. This thesis analyzes the results of the NCAA basketball tournament since 1985 and proposes several models to capture the winning seed distribution in each round. The Exponential Model estimates the winning probability of each team by modeling the time between a team&#x27;s successive winnings in a round as an exponential random variable. The Exponential Model estimates a zero probability for events that have not occurred in the training data set. The Markov Model solves this limitation by defining a Markov chain that incorporates each team&#x27;s winnings in prior rounds to estimate its winning probability. Results of these two models are validated using a chi-squared goodness of fit test. The Power Model, which is an intelligent tool for generating brackets of winners, quantifies the relative strength of each match-up in a round as a power function of the teams&#x27; seed numbers, with the exponent estimated using the historical results. The main problem of the Power Model is the data complications that are generally caused by the small size of the training data set, especially in later rounds. The Position and Upset Models solve this problem by representing the tournament&#x27;s games as a binary sequence and estimating the outcome of each game based on the teams&#x27; performance in the similar game. While generating a bracket in a forward direction from the first to the last round propagates the incorrect picks through the tournament, correctly picking the winners in later rounds automatically fills the bracket for several games in earlier rounds. This motivates developing bidirectional models that pick the winners based on a combination of models in forward and backward directions. The Power, Position, Upset, and bidirectional models are assessed based on the aggregate performance of millions of brackets for the five most recent tournaments (2012-2016). The proposed models allow one to estimate the likelihoods of different seed combinations by applying the estimated winning seed distributions, which accurately summarize the seeds&#x27; aggregate performance and provide a deeper understanding of the uncertainty in the games&#x27; outcomes.","abstract_has_math":false,"creators":["Khatibi, Arash"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":["Jacobson, Sheldon H."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2017,"date_issued":"2017-03-01T15:49:02Z","date_published":"2017-03-01T15:49:02Z","updated_at":"2026-07-22T22:26:37Z","subjects":["Bracket Challenge","NCAA Tournament"],"languages":["en"],"rights":["Copyright 2016 Arash Khatibi"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/95338","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Jacobson, Sheldon H."]},{"key":"dc:creator","label":"Author","values":["Khatibi, Arash"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2017-03-01T15:49:02Z","2016-11-23","2016-12"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Bracket Challenge","NCAA Tournament"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2016 Arash Khatibi"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/95338"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The National Collegiate Athletic Association's (NCAA) men's division I college basketball tournament is an annual competition that draws widespread attention in the United States. Estimating the outcome of each game is a popular activity undertaken by numerous websites, fans, and more recently, academic researchers. There has been a surge of interest in proposing mathematical methods to model the tournament's results and pick the winners of future games. This thesis analyzes the results of the NCAA basketball tournament since 1985 and proposes several models to capture the winning seed distribution in each round. The Exponential Model estimates the winning probability of each team by modeling the time between a team's successive winnings in a round as an exponential random variable. The Exponential Model estimates a zero probability for events that have not occurred in the training data set. The Markov Model solves this limitation by defining a Markov chain that incorporates each team's winnings in prior rounds to estimate its winning probability. Results of these two models are validated using a chi-squared goodness of fit test. The Power Model, which is an intelligent tool for generating brackets of winners, quantifies the relative strength of each match-up in a round as a power function of the teams' seed numbers, with the exponent estimated using the historical results. The main problem of the Power Model is the data complications that are generally caused by the small size of the training data set, especially in later rounds. The Position and Upset Models solve this problem by representing the tournament's games as a binary sequence and estimating the outcome of each game based on the teams' performance in the similar game. While generating a bracket in a forward direction from the first to the last round propagates the incorrect picks through the tournament, correctly picking the winners in later rounds automatically fills the bracket for several games in earlier rounds. This motivates developing bidirectional models that pick the winners based on a combination of models in forward and backward directions. The Power, Position, Upset, and bidirectional models are assessed based on the aggregate performance of millions of brackets for the five most recent tournaments (2012-2016). The proposed models allow one to estimate the likelihoods of different seed combinations by applying the estimated winning seed distributions, which accurately summarize the seeds' aggregate performance and provide a deeper understanding of the uncertainty in the games' outcomes.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-02-28 without embargo terms","The student, Arash Khatibi, accepted the attached license on 2016-11-22 at 10:42.","The student, Arash Khatibi, submitted this Thesis for approval on 2016-11-22 at 10:49.","This Thesis was approved for publication on 2016-11-23 at 09:58.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10295 on 2017-02-28 at 14:51:46","Made available in DSpace on 2017-03-01T15:49:02Z (GMT). No. of bitstreams: 2 KHATIBI-THESIS-2016.pdf: 251783 bytes, checksum: f79936220890914e8bac1f11fc370d01 (MD5) LICENSE.txt: 4210 bytes, checksum: 0008c829a88649141a17ec03b51b6684 (MD5) Previous issue date: 2016-11-23"]},{"key":"dc:format","label":"Dc Format","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Modeling the winning seed distribution of the NCAA basketball tournament"]}]}],"canonical_facts":{"dc:contributor":["Jacobson, Sheldon H."],"dc:creator":["Khatibi, Arash"],"dc:date":["2017-03-01T15:49:02Z","2016-11-23","2016-12"],"dc:description":["The National Collegiate Athletic Association's (NCAA) men's division I college basketball tournament is an annual competition that draws widespread attention in the United States. Estimating the outcome of each game is a popular activity undertaken by numerous websites, fans, and more recently, academic researchers. There has been a surge of interest in proposing mathematical methods to model the tournament's results and pick the winners of future games. This thesis analyzes the results of the NCAA basketball tournament since 1985 and proposes several models to capture the winning seed distribution in each round. The Exponential Model estimates the winning probability of each team by modeling the time between a team's successive winnings in a round as an exponential random variable. The Exponential Model estimates a zero probability for events that have not occurred in the training data set. The Markov Model solves this limitation by defining a Markov chain that incorporates each team's winnings in prior rounds to estimate its winning probability. Results of these two models are validated using a chi-squared goodness of fit test. The Power Model, which is an intelligent tool for generating brackets of winners, quantifies the relative strength of each match-up in a round as a power function of the teams' seed numbers, with the exponent estimated using the historical results. The main problem of the Power Model is the data complications that are generally caused by the small size of the training data set, especially in later rounds. The Position and Upset Models solve this problem by representing the tournament's games as a binary sequence and estimating the outcome of each game based on the teams' performance in the similar game. While generating a bracket in a forward direction from the first to the last round propagates the incorrect picks through the tournament, correctly picking the winners in later rounds automatically fills the bracket for several games in earlier rounds. This motivates developing bidirectional models that pick the winners based on a combination of models in forward and backward directions. The Power, Position, Upset, and bidirectional models are assessed based on the aggregate performance of millions of brackets for the five most recent tournaments (2012-2016). The proposed models allow one to estimate the likelihoods of different seed combinations by applying the estimated winning seed distributions, which accurately summarize the seeds' aggregate performance and provide a deeper understanding of the uncertainty in the games' outcomes.","Submission original under an indefinite embargo labeled 'Open Access'. The submission was exported from vireo on 2017-02-28 without embargo terms","The student, Arash Khatibi, accepted the attached license on 2016-11-22 at 10:42.","The student, Arash Khatibi, submitted this Thesis for approval on 2016-11-22 at 10:49.","This Thesis was approved for publication on 2016-11-23 at 09:58.","DSpace SAF Submission Ingestion Package generated from Vireo submission #10295 on 2017-02-28 at 14:51:46","Made available in DSpace on 2017-03-01T15:49:02Z (GMT). No. of bitstreams: 2 KHATIBI-THESIS-2016.pdf: 251783 bytes, checksum: f79936220890914e8bac1f11fc370d01 (MD5) LICENSE.txt: 4210 bytes, checksum: 0008c829a88649141a17ec03b51b6684 (MD5) Previous issue date: 2016-11-23"],"dc:format":["application/pdf"],"dc:identifier":["http://hdl.handle.net/2142/95338"],"dc:language":["en"],"dc:rights":["Copyright 2016 Arash Khatibi"],"dc:subject":["Bracket Challenge","NCAA Tournament"],"dc:title":["Modeling the winning seed distribution of the NCAA basketball tournament"],"dc:type":["text"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:37Z"}