Abstract
dc:description.abstractWe propose a novel principle of statistical inference motivated by information theory. In particular, we argue that if the statistician believes, as we do, that the essence of statistical inference is the compression of data, then the domain of statistical enquiry is fundamentally limited to those probability distributions $p$ with finite expected logarithms: $\log X$ must be an integrable random variable with respect to $p$. We call the set of all such distributions $\mathcal{E}$. This stands in stark contrast to what might be expected, namely a principle of finite entropy, i.e. the integrability of $\log \frac{1}{p(X)}$. We prove that $\mathcal{E}$ and the set of finite entropy distributions, $\mathcal{H}$, are in fact identical when a weak condition is met. This condition, which we call the power-law upper bound (PLUB), is itself motivated by information-theoretic concerns. In fact, when PLUB holds, another interesting discovery is made: $\mathcal{E} = \mathcal{K} = \mathcal{H}$, where $\mathcal{K}$ denotes the set of distributions with finite mean Kolmogorov complexity. In addition, we prove that this result cannot be strengthened; thus, the PLUB condition is, in a precise sense, optimal. If $X$ is a finite binary string, then $\log N(X)$ is essentially its length, where $N(X)$ denotes the positive integer corresponding to $X$ under the standard enumeration of $\mathbb{N}$ (strings ordered lexicographically by length). Understood in this way, $\log N$ is a measure of information: it is the literal number of 0s and 1s needed to express $N$ (or $X$). This places it alongside the complexity $K(N)$ and the Shannon information $\log \frac{1}{p(N)}$ as a measure of the information content of $N$; however, although it is less fundamental, $\log N$ is more practical, as $K$ is a noncomputable function and $p$ is in general unknown. By considering the set \mathcal{B}H(q) of distributions with finite cross entropy with respect to a model $q$, we then show that $\mathcal{E}$ is, in a precise sense, the largest set of distributions from which data can be compressed, i.e. for which there exists a universal source code, thus limiting the domain of information theory. *En route* we give two simple proofs that there does not exist a universal source code for $\mathcal{H}$ and we show how to significantly simplify the well known proof of [109]. We also construct a simple and explicit universal source code for $\mathcal{E}$. Having established the centrality of $\mathcal{E}$, we turn to applications, where we show that the universal assumption "$p \in \mathcal{E}$" resolves a number of difficulties that have heretofore vexed maximum-entropy inference (MaxEnt). We also argue that this principle answers many of the criticisms of MaxEnt by those who advocate for maximizing alternative "entropies." We argue that "$p \in \mathcal{E}$" leads naturally to a method for constructing robust Bayesian priors that still allows for a large degree of subjective input, thus, in a sense, bridging the divide between subjective and objective Bayesianism.
Degree
thesis:*- Name dc:type.qualificationname
- Doctor of Philosophy (PhD)
- Level dc:type.qualificationlevel
- Doctoral
- Grantor dc:publisher.institution
- University of Cambridge
- Year dc:date.issued
- 2023
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- McKittrick, Jordan
- Advisor dc:contributor.advisor
-
- Knowles, Tuomas
Subjects
dc:subject × 10Rights
dc:rightsIdentifiers
dc:identifier.*- DOI dc:identifier.doi
- https://doi.org/10.17863/CAM.111540
- OAI identifier oai:identifier
- oai:www.repository.cam.ac.uk:1810/372887