{"id":{"repo_id":"aachen","oai_identifier":"oai:publications.rwth-aachen.de:59246"},"canonical_url":"https://search.dev.ndltd.org/etd/aachen/oai:publications.rwth-aachen.de:59246","repository":{"repo_id":"aachen","name":"RWTH Aachen University","base_url":"https://publications.rwth-aachen.de/oai2d"},"display":{"title":"Architectures enabling scalable Internet search","abstract":"The vast amount of Internet content becomes manageable mainly by means of search engines that allow users to enter queries into a web form and receive as result a list of matches that refer to Intenet content elements, such as the URLs identifying matching HTML pages. However, the quality of these search engines suffers from two conceptual problems. The content volume grows faster than the bandwidth available to index it, and a large and growing share is ``hidden'' in the {em deep web}, e.g. behind HTML forms, making it hard to reach and index by search engines. The work presented here shows that these problems can be overcome if the paradigm of Internet search is reversed: content providers have to assist in making their content searchable. This leads to a distributed architecture that scales better than the central approach that current search engines implement, and that makes the deep web searchable. A UML model of the distributed search architecture was created and then implemented using Java, verifying the feasibility of the concepts. The scalability of the solution was proven using a formal model of the bandwidth consumed by a specific class of distributed search algorithms, as used by the suggested architecture. The remaining problem of how to create the content so that it complies with the suggested search architecture was tackled in two ways. Adapters for existing content can be created with little effort, as has been shown by a prototype. New Internet applications can be made searchable using the Model-Driven Architecture approach as introduced by the Object Management Group. A metamodel with a corresponding UML profile was defined that allows for a compact specification of an application's searchability. Using model transformations, a large share of the code that implements the specified searchability can be generated automatically from the models expressed in this metamodel.","abstract_html":"The vast amount of Internet content becomes manageable mainly by means of search engines that allow users to enter queries into a web form and receive as result a list of matches that refer to Intenet content elements, such as the URLs identifying matching HTML pages. However, the quality of these search engines suffers from two conceptual problems. The content volume grows faster than the bandwidth available to index it, and a large and growing share is ``hidden&#x27;&#x27; in the {em deep web}, e.g. behind HTML forms, making it hard to reach and index by search engines. The work presented here shows that these problems can be overcome if the paradigm of Internet search is reversed: content providers have to assist in making their content searchable. This leads to a distributed architecture that scales better than the central approach that current search engines implement, and that makes the deep web searchable. A UML model of the distributed search architecture was created and then implemented using Java, verifying the feasibility of the concepts. The scalability of the solution was proven using a formal model of the bandwidth consumed by a specific class of distributed search algorithms, as used by the suggested architecture. The remaining problem of how to create the content so that it complies with the suggested search architecture was tackled in two ways. Adapters for existing content can be created with little effort, as has been shown by a prototype. New Internet applications can be made searchable using the Model-Driven Architecture approach as introduced by the Object Management Group. A metamodel with a corresponding UML profile was defined that allows for a compact specification of an application&#x27;s searchability. Using model transformations, a large share of the code that implements the specified searchability can be generated automatically from the models expressed in this metamodel.","abstract_has_math":false,"creators":["Uhl, Axel"],"institution":"Publikationsserver der RWTH Aachen University","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":null,"school":null,"contributors":["Lichter, Horst"],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2004,"date_issued":"2004","date_published":"2004","updated_at":"2026-07-30T19:42:39Z","subjects":["info:eu-repo/classification/ddc/004","Informationssystem","Ereignisgesteuertes System","Integration","Benachrichtigungsdienst","Benutzerorientierung","Informatik","Internet Search","Modeling","Model-Driven Architecture","Bandwidth Model","UML","Search Infrastructures"],"languages":["eng"],"rights":["info:eu-repo/semantics/openAccess"],"rights_urls":[],"identifier_entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121050%22"],"render_values":[{"text":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121050%22","href":"https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121050%22","code":true}]}]},"links":{"outbound_url":"https://publications.rwth-aachen.de/record/59246","outbound_label":"Repository record","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Lichter, Horst"]},{"key":"dc:creator","label":"Author","values":["Uhl, Axel"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:coverage","label":"Dc Coverage","values":["DE"]},{"key":"dc:date","label":"Dc Date","values":["2004"]},{"key":"dc:publisher","label":"Institution","values":["Publikationsserver der RWTH Aachen University"]},{"key":"dc:relation","label":"Dc Relation","values":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-7273"]},{"key":"dc:type","label":"Dc Type","values":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["info:eu-repo/classification/ddc/004","Informationssystem","Ereignisgesteuertes System","Integration","Benachrichtigungsdienst","Benutzerorientierung","Informatik","Internet Search","Modeling","Model-Driven Architecture","Bandwidth Model","UML","Search Infrastructures"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["info:eu-repo/semantics/openAccess"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["https://publications.rwth-aachen.de/record/59246","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121050%22"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["The vast amount of Internet content becomes manageable mainly by means of search engines that allow users to enter queries into a web form and receive as result a list of matches that refer to Intenet content elements, such as the URLs identifying matching HTML pages. However, the quality of these search engines suffers from two conceptual problems. The content volume grows faster than the bandwidth available to index it, and a large and growing share is ``hidden'' in the {em deep web}, e.g. behind HTML forms, making it hard to reach and index by search engines. The work presented here shows that these problems can be overcome if the paradigm of Internet search is reversed: content providers have to assist in making their content searchable. This leads to a distributed architecture that scales better than the central approach that current search engines implement, and that makes the deep web searchable. A UML model of the distributed search architecture was created and then implemented using Java, verifying the feasibility of the concepts. The scalability of the solution was proven using a formal model of the bandwidth consumed by a specific class of distributed search algorithms, as used by the suggested architecture. The remaining problem of how to create the content so that it complies with the suggested search architecture was tackled in two ways. Adapters for existing content can be created with little effort, as has been shown by a prototype. New Internet applications can be made searchable using the Model-Driven Architecture approach as introduced by the Object Management Group. A metamodel with a corresponding UML profile was defined that allows for a compact specification of an application's searchability. Using model transformations, a large share of the code that implements the specified searchability can be generated automatically from the models expressed in this metamodel."]},{"key":"dc:source","label":"Dc Source","values":["Aachen : Publikationsserver der RWTH Aachen University X, 144 S. : graph. Darst. (2004). = Aachen, Techn. Hochsch., Diss., 2003"]},{"key":"dc:title","label":"Title","values":["Architectures enabling scalable Internet search"]}]}],"canonical_facts":{"dc:contributor":["Lichter, Horst"],"dc:coverage":["DE"],"dc:creator":["Uhl, Axel"],"dc:date":["2004"],"dc:description":["The vast amount of Internet content becomes manageable mainly by means of search engines that allow users to enter queries into a web form and receive as result a list of matches that refer to Intenet content elements, such as the URLs identifying matching HTML pages. However, the quality of these search engines suffers from two conceptual problems. The content volume grows faster than the bandwidth available to index it, and a large and growing share is ``hidden'' in the {em deep web}, e.g. behind HTML forms, making it hard to reach and index by search engines. The work presented here shows that these problems can be overcome if the paradigm of Internet search is reversed: content providers have to assist in making their content searchable. This leads to a distributed architecture that scales better than the central approach that current search engines implement, and that makes the deep web searchable. A UML model of the distributed search architecture was created and then implemented using Java, verifying the feasibility of the concepts. The scalability of the solution was proven using a formal model of the bandwidth consumed by a specific class of distributed search algorithms, as used by the suggested architecture. The remaining problem of how to create the content so that it complies with the suggested search architecture was tackled in two ways. Adapters for existing content can be created with little effort, as has been shown by a prototype. New Internet applications can be made searchable using the Model-Driven Architecture approach as introduced by the Object Management Group. A metamodel with a corresponding UML profile was defined that allows for a compact specification of an application's searchability. Using model transformations, a large share of the code that implements the specified searchability can be generated automatically from the models expressed in this metamodel."],"dc:identifier":["https://publications.rwth-aachen.de/record/59246","https://publications.rwth-aachen.de/search?p=id:%22RWTH-CONV-121050%22"],"dc:language":["eng"],"dc:publisher":["Publikationsserver der RWTH Aachen University"],"dc:relation":["info:eu-repo/semantics/altIdentifier/urn/urn:nbn:de:hbz:82-opus-7273"],"dc:rights":["info:eu-repo/semantics/openAccess"],"dc:source":["Aachen : Publikationsserver der RWTH Aachen University X, 144 S. : graph. Darst. (2004). = Aachen, Techn. Hochsch., Diss., 2003"],"dc:subject":["info:eu-repo/classification/ddc/004","Informationssystem","Ereignisgesteuertes System","Integration","Benachrichtigungsdienst","Benutzerorientierung","Informatik","Internet Search","Modeling","Model-Driven Architecture","Bandwidth Model","UML","Search Infrastructures"],"dc:title":["Architectures enabling scalable Internet search"],"dc:type":["info:eu-repo/semantics/doctoralThesis","info:eu-repo/semantics/publishedVersion"]},"updated_at":"2026-07-30T19:42:39Z"}