{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/143350"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/143350","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Stable Machine Learning","abstract":"This thesis explores one of the most fundamental questions in Machine Learning, namely, how should the \"learning\" component in Machine Learning be done? For essentially the entire history of the field, ever since Mosteller and Tukey proposed the paradigm in 1968, the answer has remained constant: use randomization. Namely, randomly split your data into training, validation, and test sets, then train your model on the training set, pick parameters based on the validation set, and then report performance based on the test set. Conceptually and practically simple, this methodology has gained near unanimous adoption. Despite this popularity, however, the methodology is fraught with numerous issues relating to the instability of the trained models, and the question remains whether or not we can do better? In this thesis, we answer that question in the affirmative. By taking a robust, combinatorial optimization approach, we propose a new way of training all machine learning models based on optimization rather than randomization. Rather than requesting that the model be performant against a single, randomly chosen training set, as is typically done, instead we require that it be robust against every training set of a fixed size. In this way, we extract out that which is common amongst all training sets, rather than the idiosyncrasies of any particular dataset, which are unlikely to generalize to new, yet unseen datasets. We begin by developing the methodology within the context of spatial, cross-sectional methods, and then proceed to extend the framework to time-series methods, where the contiguous structure of time now plays a key role. We next derive efficient algorithms that make the approach extremely scalable. Finally, we demonstrate the efficacy of the methodology across all methods on a large set of datasets, synthetic and real, derived from both academia and industry.","abstract_html":"This thesis explores one of the most fundamental questions in Machine Learning, namely, how should the &quot;learning&quot; component in Machine Learning be done? For essentially the entire history of the field, ever since Mosteller and Tukey proposed the paradigm in 1968, the answer has remained constant: use randomization. Namely, randomly split your data into training, validation, and test sets, then train your model on the training set, pick parameters based on the validation set, and then report performance based on the test set. Conceptually and practically simple, this methodology has gained near unanimous adoption. Despite this popularity, however, the methodology is fraught with numerous issues relating to the instability of the trained models, and the question remains whether or not we can do better? In this thesis, we answer that question in the affirmative. By taking a robust, combinatorial optimization approach, we propose a new way of training all machine learning models based on optimization rather than randomization. Rather than requesting that the model be performant against a single, randomly chosen training set, as is typically done, instead we require that it be robust against every training set of a fixed size. In this way, we extract out that which is common amongst all training sets, rather than the idiosyncrasies of any particular dataset, which are unlikely to generalize to new, yet unseen datasets. We begin by developing the methodology within the context of spatial, cross-sectional methods, and then proceed to extend the framework to time-series methods, where the contiguous structure of time now plays a key role. We next derive efficient algorithms that make the approach extremely scalable. Finally, we demonstrate the efficacy of the methodology across all methods on a large set of datasets, synthetic and real, derived from both academia and industry.","abstract_has_math":false,"creators":["Paskov, Ivan Spassimirov"],"institution":"Massachusetts Institute of Technology","degree_name":"Doctoral","degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Operations Research Center","school":null,"contributors":[],"advisors":["Bertsimas, Dimitris"],"committee_chairs":[],"committee_members":[],"year":2022,"date_issued":"2022-02","date_published":"2022-02","updated_at":"2026-07-22T22:21:02Z","subjects":[],"languages":[],"rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"rights_urls":["http://rightsstatements.org/page/InC-EDU/1.0/"],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1721.1/143350","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Bertsimas, Dimitris"]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Operations Research Center"]},{"key":"dc:creator","label":"Author","values":["Paskov, Ivan Spassimirov"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2022-06-15T13:14:24Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-06-15T13:14:24Z"]},{"key":"dc:date.issued","label":"Date","values":["2022-02"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctoral","Doctor of Philosophy"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["In Copyright - Educational Use Permitted","Copyright MIT"]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://rightsstatements.org/page/InC-EDU/1.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1721.1/143350"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["This thesis explores one of the most fundamental questions in Machine Learning, namely, how should the \"learning\" component in Machine Learning be done? For essentially the entire history of the field, ever since Mosteller and Tukey proposed the paradigm in 1968, the answer has remained constant: use randomization. Namely, randomly split your data into training, validation, and test sets, then train your model on the training set, pick parameters based on the validation set, and then report performance based on the test set. Conceptually and practically simple, this methodology has gained near unanimous adoption. Despite this popularity, however, the methodology is fraught with numerous issues relating to the instability of the trained models, and the question remains whether or not we can do better? In this thesis, we answer that question in the affirmative. By taking a robust, combinatorial optimization approach, we propose a new way of training all machine learning models based on optimization rather than randomization. Rather than requesting that the model be performant against a single, randomly chosen training set, as is typically done, instead we require that it be robust against every training set of a fixed size. In this way, we extract out that which is common amongst all training sets, rather than the idiosyncrasies of any particular dataset, which are unlikely to generalize to new, yet unseen datasets. We begin by developing the methodology within the context of spatial, cross-sectional methods, and then proceed to extend the framework to time-series methods, where the contiguous structure of time now plays a key role. We next derive efficient algorithms that make the approach extremely scalable. Finally, we demonstrate the efficacy of the methodology across all methods on a large set of datasets, synthetic and real, derived from both academia and industry."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["Ph.D."]},{"key":"dc:title","label":"Title","values":["Stable Machine Learning"]}]}],"canonical_facts":{"dc:contributor.advisor":["Bertsimas, Dimitris"],"dc:contributor.department":["Massachusetts Institute of Technology. Operations Research Center"],"dc:creator":["Paskov, Ivan Spassimirov"],"dc:date.accessioned":["2022-06-15T13:14:24Z"],"dc:date.available":["2022-06-15T13:14:24Z"],"dc:date.issued":["2022-02"],"dc:description.abstract":["This thesis explores one of the most fundamental questions in Machine Learning, namely, how should the \"learning\" component in Machine Learning be done? For essentially the entire history of the field, ever since Mosteller and Tukey proposed the paradigm in 1968, the answer has remained constant: use randomization. Namely, randomly split your data into training, validation, and test sets, then train your model on the training set, pick parameters based on the validation set, and then report performance based on the test set. Conceptually and practically simple, this methodology has gained near unanimous adoption. Despite this popularity, however, the methodology is fraught with numerous issues relating to the instability of the trained models, and the question remains whether or not we can do better? In this thesis, we answer that question in the affirmative. By taking a robust, combinatorial optimization approach, we propose a new way of training all machine learning models based on optimization rather than randomization. Rather than requesting that the model be performant against a single, randomly chosen training set, as is typically done, instead we require that it be robust against every training set of a fixed size. In this way, we extract out that which is common amongst all training sets, rather than the idiosyncrasies of any particular dataset, which are unlikely to generalize to new, yet unseen datasets. We begin by developing the methodology within the context of spatial, cross-sectional methods, and then proceed to extend the framework to time-series methods, where the contiguous structure of time now plays a key role. We next derive efficient algorithms that make the approach extremely scalable. Finally, we demonstrate the efficacy of the methodology across all methods on a large set of datasets, synthetic and real, derived from both academia and industry."],"dc:description.degree":["Ph.D."],"dc:identifier.uri":["https://hdl.handle.net/1721.1/143350"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["In Copyright - Educational Use Permitted","Copyright MIT"],"dc:rights.uri":["http://rightsstatements.org/page/InC-EDU/1.0/"],"dc:title":["Stable Machine Learning"],"dc:type":["Thesis"],"thesis:degree_name":["Doctoral","Doctor of Philosophy"]},"updated_at":"2026-07-22T22:21:02Z"}