{"id":{"repo_id":"bielefeld","oai_identifier":"oai:pub.uni-bielefeld.de:3002027"},"canonical_url":"https://search.dev.ndltd.org/etd/bielefeld/oai:pub.uni-bielefeld.de:3002027","repository":{"repo_id":"bielefeld","name":"Universität Bielefeld","base_url":"https://pub.uni-bielefeld.de/oai"},"display":{"title":"Multilingual Question Answering over Knowledge Graphs building on a Model of the Lexicon-ontology Interface","abstract":"As the amount of available linked data keeps growing, intuitive and effective paradigms for accessing and querying this body of knowledge become increasingly important. Most question answering over linked (QALD) are induced from pairs of questions and answers using various machine-learning techniques. These systems often suffer from a lack of controllability, which makes governing and incrementally improving them challenging, not to mention the initial effort required for collecting and providing training data. Furthermore, despite the transformation of linked data into a ’global database’ in which information is accessible across languages, English continues to dominate QALD research communities, often to the detriment of the approximately 7, 000 languages spoken worldwide.<br /> Based on preliminary research work, this thesis presents a multilingual question- answering (QA) approach (QueGG-Multi) that relies on a model of a lexicon-ontology interface in which the meaning of lexical entries is specified with respect to a given vocabulary or ontology. More specifically, given a lexicon in a particular language and for a particular vocabulary, the approach automatically generates a grammar that allows questions in the specific language to be parsed into SPARQL queries. The research shows that this approach outperforms current machine-learning-based QA approaches on QALD benchmarks. Furthermore, it demonstrates the extensibil- ity of the approach to different languages by adapting it to German, Italian, and Spanish.<br /> While the approach provides state-of-the-art performance on benchmarks, a crucial prerequisite for the grammar generation approach is the availability of a Lemon lexicon that describes which lexical entries can be used to verbalise the elements of a particular dataset in a specific language. To this end, the work shows that such a lexicon can be induced automatically to some extent using the LexExMachina approach that builds on association rules to find correspondences between lexical elements and ontological vocabulary elements. The results show that the method could be used to bootstrap the semiautomatic creation of a lexicon, thus having the potential to significantly reduce the human effort involved in producing lexicons for languages beyond English.","abstract_html":"As the amount of available linked data keeps growing, intuitive and effective paradigms for accessing and querying this body of knowledge become increasingly important. Most question answering over linked (QALD) are induced from pairs of questions and answers using various machine-learning techniques. These systems often suffer from a lack of controllability, which makes governing and incrementally improving them challenging, not to mention the initial effort required for collecting and providing training data. Furthermore, despite the transformation of linked data into a ’global database’ in which information is accessible across languages, English continues to dominate QALD research communities, often to the detriment of the approximately 7, 000 languages spoken worldwide.&lt;br /&gt; Based on preliminary research work, this thesis presents a multilingual question- answering (QA) approach (QueGG-Multi) that relies on a model of a lexicon-ontology interface in which the meaning of lexical entries is specified with respect to a given vocabulary or ontology. More specifically, given a lexicon in a particular language and for a particular vocabulary, the approach automatically generates a grammar that allows questions in the specific language to be parsed into SPARQL queries. The research shows that this approach outperforms current machine-learning-based QA approaches on QALD benchmarks. Furthermore, it demonstrates the extensibil- ity of the approach to different languages by adapting it to German, Italian, and Spanish.&lt;br /&gt; While the approach provides state-of-the-art performance on benchmarks, a crucial prerequisite for the grammar generation approach is the availability of a Lemon lexicon that describes which lexical entries can be used to verbalise the elements of a particular dataset in a specific language. To this end, the work shows that such a lexicon can be induced automatically to some extent using the LexExMachina approach that builds on association rules to find correspondences between lexical elements and ontological vocabulary elements. The results show that the method could be used to bootstrap the semiautomatic creation of a lexicon, thus having the potential to significantly reduce the human effort involved in producing lexicons for languages beyond English.","abstract_has_math":false,"creators":["Elahi, Mohammad Fazleh"],"institution":"Universität Bielefeld","degree_name":null,"degree_level":"thesis.doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2024,"date_issued":"2024-05-14","date_published":"2024-05-14","updated_at":"2026-07-27T18:50:04Z","subjects":[],"languages":[],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://pub.uni-bielefeld.de/record/3002027","outbound_label":"Repository record","outbound_source":"source_url"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Elahi, Mohammad Fazleh"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:publisher","label":"Institution","values":["Universitätsbibliothek Bielefeld"]},{"key":"dc:type","label":"Dc Type","values":["doctoralThesis"]},{"key":"thesis:degree_level","label":"Degree Level","values":["thesis.doctoral"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Universität Bielefeld"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["As the amount of available linked data keeps growing, intuitive and effective paradigms for accessing and querying this body of knowledge become increasingly important. Most question answering over linked (QALD) are induced from pairs of questions and answers using various machine-learning techniques. These systems often suffer from a lack of controllability, which makes governing and incrementally improving them challenging, not to mention the initial effort required for collecting and providing training data. Furthermore, despite the transformation of linked data into a ’global database’ in which information is accessible across languages, English continues to dominate QALD research communities, often to the detriment of the approximately 7, 000 languages spoken worldwide.<br /> Based on preliminary research work, this thesis presents a multilingual question- answering (QA) approach (QueGG-Multi) that relies on a model of a lexicon-ontology interface in which the meaning of lexical entries is specified with respect to a given vocabulary or ontology. More specifically, given a lexicon in a particular language and for a particular vocabulary, the approach automatically generates a grammar that allows questions in the specific language to be parsed into SPARQL queries. The research shows that this approach outperforms current machine-learning-based QA approaches on QALD benchmarks. Furthermore, it demonstrates the extensibil- ity of the approach to different languages by adapting it to German, Italian, and Spanish.<br /> While the approach provides state-of-the-art performance on benchmarks, a crucial prerequisite for the grammar generation approach is the availability of a Lemon lexicon that describes which lexical entries can be used to verbalise the elements of a particular dataset in a specific language. To this end, the work shows that such a lexicon can be induced automatically to some extent using the LexExMachina approach that builds on association rules to find correspondences between lexical elements and ontological vocabulary elements. The results show that the method could be used to bootstrap the semiautomatic creation of a lexicon, thus having the potential to significantly reduce the human effort involved in producing lexicons for languages beyond English."]},{"key":"dc:format.medium","label":"Dc Format Medium","values":["application/zip"]},{"key":"dc:title","label":"Title","values":["Multilingual Question Answering over Knowledge Graphs building on a Model of the Lexicon-ontology Interface"]}]}],"canonical_facts":{"dc:creator":["Elahi, Mohammad Fazleh"],"dc:description.abstract":["As the amount of available linked data keeps growing, intuitive and effective paradigms for accessing and querying this body of knowledge become increasingly important. Most question answering over linked (QALD) are induced from pairs of questions and answers using various machine-learning techniques. These systems often suffer from a lack of controllability, which makes governing and incrementally improving them challenging, not to mention the initial effort required for collecting and providing training data. Furthermore, despite the transformation of linked data into a ’global database’ in which information is accessible across languages, English continues to dominate QALD research communities, often to the detriment of the approximately 7, 000 languages spoken worldwide.<br /> Based on preliminary research work, this thesis presents a multilingual question- answering (QA) approach (QueGG-Multi) that relies on a model of a lexicon-ontology interface in which the meaning of lexical entries is specified with respect to a given vocabulary or ontology. More specifically, given a lexicon in a particular language and for a particular vocabulary, the approach automatically generates a grammar that allows questions in the specific language to be parsed into SPARQL queries. The research shows that this approach outperforms current machine-learning-based QA approaches on QALD benchmarks. Furthermore, it demonstrates the extensibil- ity of the approach to different languages by adapting it to German, Italian, and Spanish.<br /> While the approach provides state-of-the-art performance on benchmarks, a crucial prerequisite for the grammar generation approach is the availability of a Lemon lexicon that describes which lexical entries can be used to verbalise the elements of a particular dataset in a specific language. To this end, the work shows that such a lexicon can be induced automatically to some extent using the LexExMachina approach that builds on association rules to find correspondences between lexical elements and ontological vocabulary elements. The results show that the method could be used to bootstrap the semiautomatic creation of a lexicon, thus having the potential to significantly reduce the human effort involved in producing lexicons for languages beyond English."],"dc:format.medium":["application/zip"],"dc:publisher":["Universitätsbibliothek Bielefeld"],"dc:title":["Multilingual Question Answering over Knowledge Graphs building on a Model of the Lexicon-ontology Interface"],"dc:type":["doctoralThesis"],"thesis:degree_level":["thesis.doctoral"],"thesis:institution_name":["Universität Bielefeld"]},"updated_at":"2026-07-27T18:50:04Z"}