{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/70725"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/70725","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"Outliers in a Linear Regression Model","abstract":"Over the last several decades the linear regression model has become one of the most widely used tools of the social sciences and the physical sciences. Given the data, the least squares method gives information for statistical inferences. However, the researcher frequently feels that the regression results are not trustworthy because of possible problems with the data. These problems have sometimes been ignored in practice. It is absurd that we include all data without question if some of the data are in error, or they come from a different regime. Those data are called outliers and should be excluded from the sample or at least treated carefully.","abstract_html":"Over the last several decades the linear regression model has become one of the most widely used tools of the social sciences and the physical sciences. Given the data, the least squares method gives information for statistical inferences. However, the researcher frequently feels that the regression results are not trustworthy because of possible problems with the data. These problems have sometimes been ignored in practice. It is absurd that we include all data without question if some of the data are in error, or they come from a different regime. Those data are called outliers and should be excluded from the sample or at least treated carefully.","abstract_has_math":false,"creators":["Miyashita, Hiroshi"],"institution":"University of Illinois at Urbana-Champaign","degree_name":"Ph.D.","degree_level":"Dissertation","degree_discipline":"Economics","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-12-16T04:04:46Z","date_published":"2014-12-16T04:04:46Z","updated_at":"2026-07-22T22:26:03Z","subjects":["Statistics"],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["(UMI)AAI8218524"],"render_values":[{"text":"(UMI)AAI8218524","href":null,"code":true}]}]},"links":{"outbound_url":"http://hdl.handle.net/2142/70725","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Miyashita, Hiroshi"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-12-16T04:04:46Z","10000-01-01","1982"]},{"key":"dc:type","label":"Dc Type","values":["text"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Economics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Ph.D."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Statistics"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/70725","(UMI)AAI8218524"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Over the last several decades the linear regression model has become one of the most widely used tools of the social sciences and the physical sciences. Given the data, the least squares method gives information for statistical inferences. However, the researcher frequently feels that the regression results are not trustworthy because of possible problems with the data. These problems have sometimes been ignored in practice. It is absurd that we include all data without question if some of the data are in error, or they come from a different regime. Those data are called outliers and should be excluded from the sample or at least treated carefully.","Several test statistics for detecting outliers have been developed. However, the tests based on those statistics usually require the assumption of error normality. If the underlying error distribution deviates from the normal, the test is not trustworthy. This is confirmed by a simulation study. Even if the error distribution is normal, it is computationally burdensome and sometimes impossible to locate more than one outlier correctly. Therefore, it is impossible to detect outliers if the error distribution is non-normal or there is a possibility of having more than one outlier. One solution to this problem is a Bayesian approach.","The test of significance has little relevance in the context of a Bayesian approach. We accommodate outliers rather than detect and drop them. Furthermore, we don't have to know how many outliers exist in the sample. By constructing an appropriate prior distribution of having outliers, we can derive a posterior distribution of the regression parameters. The underlying error distribution is not restricted to the normal. Introducing a class of symmetric exponential power distributions which includes the normal as a special case, we can handle the situation in which the error distribution is assumed to be non-normal. Hypothesis testing can be done by constructing a Bayesian confidence interval. Using the interval we can test a null hypothesis in the possible presence of outliers.","Made available in DSpace on 2014-12-16T04:04:46Z (GMT). No. of bitstreams: 1 8218524.pdf: 6767396 bytes, checksum: 7af1f8b6d12c7f2e3d1bd39755d2bd61 (MD5) Previous issue date: 1982","Embargo set by: Seth Robbins for item 70891 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","228 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 1982."]},{"key":"dc:title","label":"Title","values":["Outliers in a Linear Regression Model"]}]}],"canonical_facts":{"dc:creator":["Miyashita, Hiroshi"],"dc:date":["2014-12-16T04:04:46Z","10000-01-01","1982"],"dc:description":["Over the last several decades the linear regression model has become one of the most widely used tools of the social sciences and the physical sciences. Given the data, the least squares method gives information for statistical inferences. However, the researcher frequently feels that the regression results are not trustworthy because of possible problems with the data. These problems have sometimes been ignored in practice. It is absurd that we include all data without question if some of the data are in error, or they come from a different regime. Those data are called outliers and should be excluded from the sample or at least treated carefully.","Several test statistics for detecting outliers have been developed. However, the tests based on those statistics usually require the assumption of error normality. If the underlying error distribution deviates from the normal, the test is not trustworthy. This is confirmed by a simulation study. Even if the error distribution is normal, it is computationally burdensome and sometimes impossible to locate more than one outlier correctly. Therefore, it is impossible to detect outliers if the error distribution is non-normal or there is a possibility of having more than one outlier. One solution to this problem is a Bayesian approach.","The test of significance has little relevance in the context of a Bayesian approach. We accommodate outliers rather than detect and drop them. Furthermore, we don't have to know how many outliers exist in the sample. By constructing an appropriate prior distribution of having outliers, we can derive a posterior distribution of the regression parameters. The underlying error distribution is not restricted to the normal. Introducing a class of symmetric exponential power distributions which includes the normal as a special case, we can handle the situation in which the error distribution is assumed to be non-normal. Hypothesis testing can be done by constructing a Bayesian confidence interval. Using the interval we can test a null hypothesis in the possible presence of outliers.","Made available in DSpace on 2014-12-16T04:04:46Z (GMT). No. of bitstreams: 1 8218524.pdf: 6767396 bytes, checksum: 7af1f8b6d12c7f2e3d1bd39755d2bd61 (MD5) Previous issue date: 1982","Embargo set by: Seth Robbins for item 70891 Lift date: Forever Reason: Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","Restricted to the U of I community idenfinitely during batch ingest of legacy ETDs","U of I Only","228 p.","Thesis (Ph.D.)--University of Illinois at Urbana-Champaign, 1982."],"dc:identifier":["http://hdl.handle.net/2142/70725","(UMI)AAI8218524"],"dc:subject":["Statistics"],"dc:title":["Outliers in a Linear Regression Model"],"dc:type":["text"],"thesis:degree_discipline":["Economics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Ph.D."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:26:03Z"}