{"id":{"repo_id":"lsu-thes","oai_identifier":"oai:repository.lsu.edu:gradschool_dissertations-1784"},"canonical_url":"https://search.dev.ndltd.org/etd/lsu-thes/oai:repository.lsu.edu:gradschool_dissertations-1784","repository":{"repo_id":"lsu-thes","name":"Lousiana State University","base_url":"https://repository.lsu.edu/do/oai/"},"display":{"title":"Efficient Indexing for Structured and Unstructured Data","abstract":"The collection of digital data is growing at an exponential rate. Data originates from wide range of data sources such as text feeds, biological sequencers, internet traffic over routers, through sensors and many other sources. To mine intelligent information from these sources, users have to query the data. Indexing techniques aim to reduce the query time by preprocessing the data. Diversity of data sources in real world makes it imperative to develop application specific indexing solutions based on the data to be queried. Data can be structured i.e., relational tables or unstructured i.e., free text. Moreover, increasingly many applications need to seamlessly analyze both kinds of data making data integration a central issue. Integrating text with structured data needs to account for missing values, errors in the data etc. Probabilistic models have been proposed recently for this purpose. These models are also useful for applications where uncertainty is inherent in data e.g. sensor networks. This dissertation aims to propose efficient indexing solutions for several problems that lie at the intersection of database and information retrieval such as joining ranked inputs, full-text documents searching etc. Other well-known problems of ranked retrieval and pattern matching are also studied under probabilistic settings. For each problem, the worst-case theoretical bounds of the proposed solutions are established and/or their practicality is demonstrated by thorough experimentation.","abstract_html":"The collection of digital data is growing at an exponential rate. Data originates from wide range of data sources such as text feeds, biological sequencers, internet traffic over routers, through sensors and many other sources. To mine intelligent information from these sources, users have to query the data. Indexing techniques aim to reduce the query time by preprocessing the data. Diversity of data sources in real world makes it imperative to develop application specific indexing solutions based on the data to be queried. Data can be structured i.e., relational tables or unstructured i.e., free text. Moreover, increasingly many applications need to seamlessly analyze both kinds of data making data integration a central issue. Integrating text with structured data needs to account for missing values, errors in the data etc. Probabilistic models have been proposed recently for this purpose. These models are also useful for applications where uncertainty is inherent in data e.g. sensor networks. This dissertation aims to propose efficient indexing solutions for several problems that lie at the intersection of database and information retrieval such as joining ranked inputs, full-text documents searching etc. Other well-known problems of ranked retrieval and pattern matching are also studied under probabilistic settings. For each problem, the worst-case theoretical bounds of the proposed solutions are established and/or their practicality is demonstrated by thorough experimentation.","abstract_has_math":false,"creators":["Patil, Manish Madhukar"],"institution":"Computer Science","degree_name":"Doctor of Philosophy (PhD)","degree_level":"Dissertation","degree_discipline":"Computer Sciences","degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2014,"date_issued":"2014-01-01T08:00:00Z","date_published":"2014-01-01T08:00:00Z","updated_at":"2026-07-24T02:58:17Z","subjects":["Indexing","Top-k","Document Retrieval","Probabilistic Data","Uncertain Data"],"languages":[],"rights":["unrestricted","Release the entire work immediately for access worldwide."],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["etd-08182014-125357","https://repository.lsu.edu/gradschool_dissertations/785"],"render_values":[{"text":"etd-08182014-125357","href":null,"code":true},{"text":"https://repository.lsu.edu/gradschool_dissertations/785","href":"https://repository.lsu.edu/gradschool_dissertations/785","code":true}]}]},"links":{"outbound_url":"https://doi.org/10.31390/gradschool_dissertations.785","outbound_label":"DOI","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Patil, Manish Madhukar"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2014-03-10"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2022-05-12T23:09:59Z"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Sciences"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Dissertation"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Doctor of Philosophy (PhD)"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Computer Science"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Indexing","Top-k","Document Retrieval","Probabilistic Data","Uncertain Data"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["unrestricted","Release the entire work immediately for access worldwide."]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["etd-08182014-125357","10.31390/gradschool_dissertations.785","https://repository.lsu.edu/gradschool_dissertations/785"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["The collection of digital data is growing at an exponential rate. Data originates from wide range of data sources such as text feeds, biological sequencers, internet traffic over routers, through sensors and many other sources. To mine intelligent information from these sources, users have to query the data. Indexing techniques aim to reduce the query time by preprocessing the data. Diversity of data sources in real world makes it imperative to develop application specific indexing solutions based on the data to be queried. Data can be structured i.e., relational tables or unstructured i.e., free text. Moreover, increasingly many applications need to seamlessly analyze both kinds of data making data integration a central issue. Integrating text with structured data needs to account for missing values, errors in the data etc. Probabilistic models have been proposed recently for this purpose. These models are also useful for applications where uncertainty is inherent in data e.g. sensor networks. This dissertation aims to propose efficient indexing solutions for several problems that lie at the intersection of database and information retrieval such as joining ranked inputs, full-text documents searching etc. Other well-known problems of ranked retrieval and pattern matching are also studied under probabilistic settings. For each problem, the worst-case theoretical bounds of the proposed solutions are established and/or their practicality is demonstrated by thorough experimentation."]},{"key":"dc:title","label":"Title","values":["Efficient Indexing for Structured and Unstructured Data"]}]}],"canonical_facts":{"dc:creator":["Patil, Manish Madhukar"],"dc:date":["2014-03-10"],"dc:date.available":["2022-05-12T23:09:59Z"],"dc:description.abstract":["The collection of digital data is growing at an exponential rate. Data originates from wide range of data sources such as text feeds, biological sequencers, internet traffic over routers, through sensors and many other sources. To mine intelligent information from these sources, users have to query the data. Indexing techniques aim to reduce the query time by preprocessing the data. Diversity of data sources in real world makes it imperative to develop application specific indexing solutions based on the data to be queried. Data can be structured i.e., relational tables or unstructured i.e., free text. Moreover, increasingly many applications need to seamlessly analyze both kinds of data making data integration a central issue. Integrating text with structured data needs to account for missing values, errors in the data etc. Probabilistic models have been proposed recently for this purpose. These models are also useful for applications where uncertainty is inherent in data e.g. sensor networks. This dissertation aims to propose efficient indexing solutions for several problems that lie at the intersection of database and information retrieval such as joining ranked inputs, full-text documents searching etc. Other well-known problems of ranked retrieval and pattern matching are also studied under probabilistic settings. For each problem, the worst-case theoretical bounds of the proposed solutions are established and/or their practicality is demonstrated by thorough experimentation."],"dc:identifier":["etd-08182014-125357","10.31390/gradschool_dissertations.785","https://repository.lsu.edu/gradschool_dissertations/785"],"dc:rights":["unrestricted","Release the entire work immediately for access worldwide."],"dc:subject":["Indexing","Top-k","Document Retrieval","Probabilistic Data","Uncertain Data"],"dc:title":["Efficient Indexing for Structured and Unstructured Data"],"thesis:degree_discipline":["Computer Sciences"],"thesis:degree_level":["Dissertation"],"thesis:degree_name":["Doctor of Philosophy (PhD)"],"thesis:institution_name":["Computer Science"]},"updated_at":"2026-07-24T02:58:17Z"}