{"id":{"repo_id":"toronto-retro","oai_identifier":"oai:utoronto.scholaris.ca:1807/29890"},"canonical_url":"https://search.dev.ndltd.org/etd/toronto-retro/oai:utoronto.scholaris.ca:1807/29890","repository":{"repo_id":"toronto-retro","name":"University of Toronto","base_url":"https://utoronto.scholaris.ca/server/oai/request"},"display":{"title":"Statistical Methods for Dating Collections of Historical Documents","abstract":"The problem in this thesis was originally motivated by problems presented with documents of Early England Data Set (DEEDS). The central problem with these medieval documents is the lack of methods to assign accurate dates to those documents which bear no date. With the problems of the DEEDS documents in mind, we present two methods to impute missing features of texts. In the first method, we suggest a new class of metrics for measuring distances between texts. We then show how to combine the distances between the texts using statistical smoothing. This method can be adapted to settings where the features of the texts are ordered or unordered categoricals (as in the case of, for example, authorship assignment problems). In the second method, we estimate the probability of occurrences of words in texts using nonparametric regression techniques of local polynomial fitting with kernel weight to generalized linear models. We combine the estimated probability of occurrences of words of a text to estimate the probability of occurrence of a text as a function of its feature -- the feature in this case being the date in which the text is written. The application and results of our methods to the DEEDS documents are presented.","abstract_html":"The problem in this thesis was originally motivated by problems presented with documents of Early England Data Set (DEEDS). The central problem with these medieval documents is the lack of methods to assign accurate dates to those documents which bear no date. With the problems of the DEEDS documents in mind, we present two methods to impute missing features of texts. In the first method, we suggest a new class of metrics for measuring distances between texts. We then show how to combine the distances between the texts using statistical smoothing. This method can be adapted to settings where the features of the texts are ordered or unordered categoricals (as in the case of, for example, authorship assignment problems). In the second method, we estimate the probability of occurrences of words in texts using nonparametric regression techniques of local polynomial fitting with kernel weight to generalized linear models. We combine the estimated probability of occurrences of words of a text to estimate the probability of occurrence of a text as a function of its feature -- the feature in this case being the date in which the text is written. The application and results of our methods to the DEEDS documents are presented.","abstract_has_math":false,"creators":["Tilahun, Gelila"],"institution":null,"degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Statistics","school":null,"contributors":[],"advisors":["Feuerverger, Andrey"],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-08-31","date_published":"2011-08-31","updated_at":"2026-07-27T21:28:07Z","subjects":["Kernel","Dating Documents","Shingle","Correspondence distance","Smoothing","Generalized linear models","Logistics regression","Local polynomial regression"],"languages":["en_ca"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1807/29890","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Feuerverger, Andrey"]},{"key":"dc:contributor.department","label":"Department","values":["Statistics"]},{"key":"dc:creator","label":"Author","values":["Tilahun, Gelila"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-06"]},{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2011-08-31T23:46:32Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["NO_RESTRICTION","2011-08-31T23:46:32Z"]},{"key":"dc:date.issued","label":"Date","values":["2011-08-31"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Kernel","Dating Documents","Shingle","Correspondence distance","Smoothing","Generalized linear models","Logistics regression","Local polynomial regression"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_ca"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1807/29890"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The problem in this thesis was originally motivated by problems presented with documents of Early England Data Set (DEEDS). The central problem with these medieval documents is the lack of methods to assign accurate dates to those documents which bear no date. With the problems of the DEEDS documents in mind, we present two methods to impute missing features of texts. In the first method, we suggest a new class of metrics for measuring distances between texts. We then show how to combine the distances between the texts using statistical smoothing. This method can be adapted to settings where the features of the texts are ordered or unordered categoricals (as in the case of, for example, authorship assignment problems). In the second method, we estimate the probability of occurrences of words in texts using nonparametric regression techniques of local polynomial fitting with kernel weight to generalized linear models. We combine the estimated probability of occurrences of words of a text to estimate the probability of occurrence of a text as a function of its feature -- the feature in this case being the date in which the text is written. The application and results of our methods to the DEEDS documents are presented."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["PhD"]},{"key":"dc:title","label":"Title","values":["Statistical Methods for Dating Collections of Historical Documents"]}]}],"canonical_facts":{"dc:contributor.advisor":["Feuerverger, Andrey"],"dc:contributor.department":["Statistics"],"dc:creator":["Tilahun, Gelila"],"dc:date":["2011-06"],"dc:date.accessioned":["2011-08-31T23:46:32Z"],"dc:date.available":["NO_RESTRICTION","2011-08-31T23:46:32Z"],"dc:date.issued":["2011-08-31"],"dc:description.abstract":["The problem in this thesis was originally motivated by problems presented with documents of Early England Data Set (DEEDS). The central problem with these medieval documents is the lack of methods to assign accurate dates to those documents which bear no date. With the problems of the DEEDS documents in mind, we present two methods to impute missing features of texts. In the first method, we suggest a new class of metrics for measuring distances between texts. We then show how to combine the distances between the texts using statistical smoothing. This method can be adapted to settings where the features of the texts are ordered or unordered categoricals (as in the case of, for example, authorship assignment problems). In the second method, we estimate the probability of occurrences of words in texts using nonparametric regression techniques of local polynomial fitting with kernel weight to generalized linear models. We combine the estimated probability of occurrences of words of a text to estimate the probability of occurrence of a text as a function of its feature -- the feature in this case being the date in which the text is written. The application and results of our methods to the DEEDS documents are presented."],"dc:description.degree":["PhD"],"dc:identifier.uri":["http://hdl.handle.net/1807/29890"],"dc:language.iso":["en_ca"],"dc:subject":["Kernel","Dating Documents","Shingle","Correspondence distance","Smoothing","Generalized linear models","Logistics regression","Local polynomial regression"],"dc:title":["Statistical Methods for Dating Collections of Historical Documents"],"dc:type":["Thesis"]},"updated_at":"2026-07-27T21:28:07Z"}