Back to results

University of Illinois at Urbana-Champaign

Modeling the winning seed distribution of the NCAA basketball tournament

Abstract

dc:description

The National Collegiate Athletic Association's (NCAA) men's division I college basketball tournament is an annual competition that draws widespread attention in the United States. Estimating the outcome of each game is a popular activity undertaken by numerous websites, fans, and more recently, academic researchers. There has been a surge of interest in proposing mathematical methods to model the tournament's results and pick the winners of future games. This thesis analyzes the results of the NCAA basketball tournament since 1985 and proposes several models to capture the winning seed distribution in each round. The Exponential Model estimates the winning probability of each team by modeling the time between a team's successive winnings in a round as an exponential random variable. The Exponential Model estimates a zero probability for events that have not occurred in the training data set. The Markov Model solves this limitation by defining a Markov chain that incorporates each team's winnings in prior rounds to estimate its winning probability. Results of these two models are validated using a chi-squared goodness of fit test. The Power Model, which is an intelligent tool for generating brackets of winners, quantifies the relative strength of each match-up in a round as a power function of the teams' seed numbers, with the exponent estimated using the historical results. The main problem of the Power Model is the data complications that are generally caused by the small size of the training data set, especially in later rounds. The Position and Upset Models solve this problem by representing the tournament's games as a binary sequence and estimating the outcome of each game based on the teams' performance in the similar game. While generating a bracket in a forward direction from the first to the last round propagates the incorrect picks through the tournament, correctly picking the winners in later rounds automatically fills the bracket for several games in earlier rounds. This motivates developing bidirectional models that pick the winners based on a combination of models in forward and backward directions. The Power, Position, Upset, and bidirectional models are assessed based on the aggregate performance of millions of brackets for the five most recent tournaments (2012-2016). The proposed models allow one to estimate the likelihoods of different seed combinations by applying the estimated winning seed distributions, which accurately summarize the seeds' aggregate performance and provide a deeper understanding of the uncertainty in the games' outcomes.

Degree

thesis:*
Name thesis:degree_name
M.S.
Level thesis:degree_level
Thesis
Discipline thesis:degree_discipline
Computer Science
Grantor
University of Illinois at Urbana-Champaign
Year dc:date
2017

Author and committee

dc:creator, dc:contributor.*
Author dc:creator
  • Khatibi, Arash
Contributors dc:contributor
  • Jacobson, Sheldon H.

Subjects

dc:subject × 2

Rights

dc:rights
Statement dc:rights
  • Copyright 2016 Arash Khatibi
Language dc:language
en

Identifiers

dc:identifier.*
Handle dc:identifier
http://hdl.handle.net/2142/95338
OAI identifier oai:identifier
oai:www.ideals.illinois.edu:2142/95338

Chain of custody

source
Harvested from
University of Illinois - Urbana-Champaign
Base URL
www.ideals.illinois.edu/oai-pmh
Last updated
2026-07-22
Source record
OAI-PMH GetRecord
citation

Khatibi, Arash. Modeling the winning seed distribution of the NCAA basketball tournament. Thesis thesis, University of Illinois at Urbana-Champaign, 2017. http://hdl.handle.net/2142/95338