{"id":{"repo_id":"odu","oai_identifier":"oai:digitalcommons.odu.edu:mathstat_etds-1063"},"canonical_url":"https://search.dev.ndltd.org/etd/odu/oai:digitalcommons.odu.edu:mathstat_etds-1063","repository":{"repo_id":"odu","name":"Old Dominion University","base_url":"https://digitalcommons.odu.edu/do/oai/"},"display":{"title":"Canonical Correlation and Correspondence Analysis of Longitudinal Data","abstract":"<p>Assessing the relationship between two sets of multivariate vectors is an important problem in statistics. Canonical correlation coefficients are used to study these relationships. Canonical correlation analysis (CCA) is a general multivariate method that is mainly used to study relationships when both sets of variables are quantitative. When the variables are qualitative (categorical), a technique called correspondence analysis (CA) is used. Canonical correspondence analysis (CCPA) is used to deal with the case when one set of variables is categorical and the other set is quantitative. By exploiting the interrelationships between these three techniques we first provide a theoretical basis for CCPA.</p> <p>Next, in this dissertation, we have generalized each of these three techniques to analyze the relationships between two sets of repeatedly or longitudinally observed data. When the two vectors are quantitative, we use a block Kronecker product matrix to model dependency of the variables over time. We then apply canonical correlation analysis on this matrix to obtain canonical correlations and canonical variables. When the variables are qualitative, the data are summarized in the form of a contingency table. It is generally not straightforward to model dependency of contingency tables over time. However, we have proposed fitting correlated linear models to the summary statistics obtained by performing the usual correspondence analysis at each time period. We have shown that the most useful summary measure for this purpose is the first singular value of the correspondence matrix, which is essentially the matrix of relative frequencies obtained from the given contingency table. Our method is a reasonable approach to analyze repeated contingency table data. Finally, to deal with the case when one set of variables is categorical and the other set is quantitative, we have proposed combining the two approaches to deal with quantitative and qualitative variables. We have illustrated and studied the performances of our methods my implementing them on simulated data sets.</p> <p>High dimensional data are now common due to the Internet, genomics, proteomics, and the like. Although, correspondence analysis and other methods considered in this dissertation are general techniques for analyzing multivariate data their usefulness for analyzing very high dimensional data have not been compared with the other more modern machine learning methods. In the last chapter of this dissertation, we provide a brief introduction to a machine learning method that is used to analyze very high dimensional and sparse contingency table data from the field of language processing or information retrieval, named latent semantic analysis (LSA). We then propose certain criteria to compare the performance of LSA with the correspondence analysis. Based on these criteria we find that under certain situations correspondence analysis performs better.</p>","abstract_html":"&lt;p&gt;Assessing the relationship between two sets of multivariate vectors is an important problem in statistics. Canonical correlation coefficients are used to study these relationships. Canonical correlation analysis (CCA) is a general multivariate method that is mainly used to study relationships when both sets of variables are quantitative. When the variables are qualitative (categorical), a technique called correspondence analysis (CA) is used. Canonical correspondence analysis (CCPA) is used to deal with the case when one set of variables is categorical and the other set is quantitative. By exploiting the interrelationships between these three techniques we first provide a theoretical basis for CCPA.&lt;/p&gt; &lt;p&gt;Next, in this dissertation, we have generalized each of these three techniques to analyze the relationships between two sets of repeatedly or longitudinally observed data. When the two vectors are quantitative, we use a block Kronecker product matrix to model dependency of the variables over time. We then apply canonical correlation analysis on this matrix to obtain canonical correlations and canonical variables. When the variables are qualitative, the data are summarized in the form of a contingency table. It is generally not straightforward to model dependency of contingency tables over time. However, we have proposed fitting correlated linear models to the summary statistics obtained by performing the usual correspondence analysis at each time period. We have shown that the most useful summary measure for this purpose is the first singular value of the correspondence matrix, which is essentially the matrix of relative frequencies obtained from the given contingency table. Our method is a reasonable approach to analyze repeated contingency table data. Finally, to deal with the case when one set of variables is categorical and the other set is quantitative, we have proposed combining the two approaches to deal with quantitative and qualitative variables. We have illustrated and studied the performances of our methods my implementing them on simulated data sets.&lt;/p&gt; &lt;p&gt;High dimensional data are now common due to the Internet, genomics, proteomics, and the like. Although, correspondence analysis and other methods considered in this dissertation are general techniques for analyzing multivariate data their usefulness for analyzing very high dimensional data have not been compared with the other more modern machine learning methods. In the last chapter of this dissertation, we provide a brief introduction to a machine learning method that is used to analyze very high dimensional and sparse contingency table data from the field of language processing or information retrieval, named latent semantic analysis (LSA). We then propose certain criteria to compare the performance of LSA with the correspondence analysis. Based on these criteria we find that under certain situations correspondence analysis performs better.&lt;/p&gt;","abstract_has_math":false,"creators":["Srivastava, Jayesh"],"institution":null,"degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation","degree_discipline":"Mathematics & Statistics","degree_department":null,"school":null,"contributors":["Dayanand N. Naik","N. Rao Chaganty","Larry Lee","Edward Markowski"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2007,"date_issued":"2007-04-01T07:00:00Z","date_published":"2007-04-01T07:00:00Z","updated_at":"2026-07-24T03:35:00Z","subjects":["Canonical correlation","Correspondence analysis","Applied Statistics","Longitudinal Data Analysis and Time Series"],"languages":[],"rights":["<p>In Copyright. URI: <a href=\"http://rightsstatements.org/vocab/InC/1.0/\">http://rightsstatements.org/vocab/InC/1.0/</a> This Item is protected by copyright and/or related rights. You are free to use this Item in any way that is permitted by the copyright and related rights legislation that applies to your use. For other uses you need to obtain permission from the rights-holder(s).</p>"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["9780549069508"],"render_values":[{"text":"9780549069508","href":null,"code":true}]}]},"links":{"outbound_url":"https://digitalcommons.odu.edu/mathstat_etds/65","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Dayanand N. Naik","N. Rao Chaganty","Larry Lee","Edward Markowski"]},{"key":"dc:creator","label":"Author","values":["Srivastava, Jayesh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.available","label":"Dc Date Available","values":["2019-06-13T07:00:00Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Mathematics & Statistics"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Canonical correlation","Correspondence analysis","Applied Statistics","Longitudinal Data Analysis and Time Series"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["<p>In Copyright. URI: <a href=\"http://rightsstatements.org/vocab/InC/1.0/\">http://rightsstatements.org/vocab/InC/1.0/</a> This Item is protected by copyright and/or related rights. You are free to use this Item in any way that is permitted by the copyright and related rights legislation that applies to your use. For other uses you need to obtain permission from the rights-holder(s).</p>"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["9780549069508","https://digitalcommons.odu.edu/mathstat_etds/65"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["<p>Assessing the relationship between two sets of multivariate vectors is an important problem in statistics. Canonical correlation coefficients are used to study these relationships. Canonical correlation analysis (CCA) is a general multivariate method that is mainly used to study relationships when both sets of variables are quantitative. When the variables are qualitative (categorical), a technique called correspondence analysis (CA) is used. Canonical correspondence analysis (CCPA) is used to deal with the case when one set of variables is categorical and the other set is quantitative. By exploiting the interrelationships between these three techniques we first provide a theoretical basis for CCPA.</p> <p>Next, in this dissertation, we have generalized each of these three techniques to analyze the relationships between two sets of repeatedly or longitudinally observed data. When the two vectors are quantitative, we use a block Kronecker product matrix to model dependency of the variables over time. We then apply canonical correlation analysis on this matrix to obtain canonical correlations and canonical variables. When the variables are qualitative, the data are summarized in the form of a contingency table. It is generally not straightforward to model dependency of contingency tables over time. However, we have proposed fitting correlated linear models to the summary statistics obtained by performing the usual correspondence analysis at each time period. We have shown that the most useful summary measure for this purpose is the first singular value of the correspondence matrix, which is essentially the matrix of relative frequencies obtained from the given contingency table. Our method is a reasonable approach to analyze repeated contingency table data. Finally, to deal with the case when one set of variables is categorical and the other set is quantitative, we have proposed combining the two approaches to deal with quantitative and qualitative variables. We have illustrated and studied the performances of our methods my implementing them on simulated data sets.</p> <p>High dimensional data are now common due to the Internet, genomics, proteomics, and the like. Although, correspondence analysis and other methods considered in this dissertation are general techniques for analyzing multivariate data their usefulness for analyzing very high dimensional data have not been compared with the other more modern machine learning methods. In the last chapter of this dissertation, we provide a brief introduction to a machine learning method that is used to analyze very high dimensional and sparse contingency table data from the field of language processing or information retrieval, named latent semantic analysis (LSA). We then propose certain criteria to compare the performance of LSA with the correspondence analysis. Based on these criteria we find that under certain situations correspondence analysis performs better.</p>"]},{"key":"dc:title","label":"Title","values":["Canonical Correlation and Correspondence Analysis of Longitudinal Data"]}]}],"canonical_facts":{"dc:contributor":["Dayanand N. Naik","N. Rao Chaganty","Larry Lee","Edward Markowski"],"dc:creator":["Srivastava, Jayesh"],"dc:date.available":["2019-06-13T07:00:00Z"],"dc:description.abstract":["<p>Assessing the relationship between two sets of multivariate vectors is an important problem in statistics. Canonical correlation coefficients are used to study these relationships. Canonical correlation analysis (CCA) is a general multivariate method that is mainly used to study relationships when both sets of variables are quantitative. When the variables are qualitative (categorical), a technique called correspondence analysis (CA) is used. Canonical correspondence analysis (CCPA) is used to deal with the case when one set of variables is categorical and the other set is quantitative. By exploiting the interrelationships between these three techniques we first provide a theoretical basis for CCPA.</p> <p>Next, in this dissertation, we have generalized each of these three techniques to analyze the relationships between two sets of repeatedly or longitudinally observed data. When the two vectors are quantitative, we use a block Kronecker product matrix to model dependency of the variables over time. We then apply canonical correlation analysis on this matrix to obtain canonical correlations and canonical variables. When the variables are qualitative, the data are summarized in the form of a contingency table. It is generally not straightforward to model dependency of contingency tables over time. However, we have proposed fitting correlated linear models to the summary statistics obtained by performing the usual correspondence analysis at each time period. We have shown that the most useful summary measure for this purpose is the first singular value of the correspondence matrix, which is essentially the matrix of relative frequencies obtained from the given contingency table. Our method is a reasonable approach to analyze repeated contingency table data. Finally, to deal with the case when one set of variables is categorical and the other set is quantitative, we have proposed combining the two approaches to deal with quantitative and qualitative variables. We have illustrated and studied the performances of our methods my implementing them on simulated data sets.</p> <p>High dimensional data are now common due to the Internet, genomics, proteomics, and the like. Although, correspondence analysis and other methods considered in this dissertation are general techniques for analyzing multivariate data their usefulness for analyzing very high dimensional data have not been compared with the other more modern machine learning methods. In the last chapter of this dissertation, we provide a brief introduction to a machine learning method that is used to analyze very high dimensional and sparse contingency table data from the field of language processing or information retrieval, named latent semantic analysis (LSA). We then propose certain criteria to compare the performance of LSA with the correspondence analysis. Based on these criteria we find that under certain situations correspondence analysis performs better.</p>"],"dc:identifier":["9780549069508","https://digitalcommons.odu.edu/mathstat_etds/65"],"dc:rights":["<p>In Copyright. URI: <a href=\"http://rightsstatements.org/vocab/InC/1.0/\">http://rightsstatements.org/vocab/InC/1.0/</a> This Item is protected by copyright and/or related rights. You are free to use this Item in any way that is permitted by the copyright and related rights legislation that applies to your use. For other uses you need to obtain permission from the rights-holder(s).</p>"],"dc:subject":["Canonical correlation","Correspondence analysis","Applied Statistics","Longitudinal Data Analysis and Time Series"],"dc:title":["Canonical Correlation and Correspondence Analysis of Longitudinal Data"],"thesis:degree_discipline":["Mathematics & Statistics"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy (PhD)"]},"updated_at":"2026-07-24T03:35:00Z"}