{"id":{"repo_id":"tamu","oai_identifier":"oai:oaktrust.library.tamu.edu:1969.1/ETD-TAMU-2011-12-10235"},"canonical_url":"https://search.dev.ndltd.org/etd/tamu/oai:oaktrust.library.tamu.edu:1969.1/ETD-TAMU-2011-12-10235","repository":{"repo_id":"tamu","name":"Texas A&M University","base_url":"https://oaktrust.library.tamu.edu/server/oai/request"},"display":{"title":"Identifying Search Engine Spam Using DNS","abstract":"Web crawlers encounter both finite and infinite elements during crawl. Pages and hosts can be infinitely generated using automated scripts and DNS wildcard entries. It is a challenge to rank such resources as an entire web of pages and hosts could be created to manipulate the rank of a target resource. It is crucial to be able to differentiate genuine content from spam in real-time to allocate crawl budgets. In this study, ranking algorithms to rank hosts are designed which use the finite Pay Level Domains(PLD) and IPv4 addresses. Heterogenous graphs derived from the webgraph of IRLbot are used to achieve this. PLD Supporters (PSUPP) which is the number of level-2 PLD supporters for each host on the host-host-PLD graph is the first algorithm that is studied. This is further improved by True PLD Supporters(TSUPP) which uses true egalitarian level-2 PLD supporters on the host-IP-PLD graph and DNS blacklists. It was found that support from content farms and stolen links could be eliminated by finding TSUPP. When TSUPP was applied on the host graph of IRLbot, there was less than 1% spam in the top 100,000 hosts.","abstract_html":"Web crawlers encounter both finite and infinite elements during crawl. Pages and hosts can be infinitely generated using automated scripts and DNS wildcard entries. It is a challenge to rank such resources as an entire web of pages and hosts could be created to manipulate the rank of a target resource. It is crucial to be able to differentiate genuine content from spam in real-time to allocate crawl budgets. In this study, ranking algorithms to rank hosts are designed which use the finite Pay Level Domains(PLD) and IPv4 addresses. Heterogenous graphs derived from the webgraph of IRLbot are used to achieve this. PLD Supporters (PSUPP) which is the number of level-2 PLD supporters for each host on the host-host-PLD graph is the first algorithm that is studied. This is further improved by True PLD Supporters(TSUPP) which uses true egalitarian level-2 PLD supporters on the host-IP-PLD graph and DNS blacklists. It was found that support from content farms and stolen links could be eliminated by finding TSUPP. When TSUPP was applied on the host graph of IRLbot, there was less than 1% spam in the top 100,000 hosts.","abstract_has_math":false,"creators":["Mathiharan, Siddhartha Sankaran"],"institution":"Texas A&M University","degree_name":"Master of Science","degree_level":"Masters","degree_discipline":"Computer Science","degree_department":null,"school":null,"contributors":[],"advisors":["Loguinov, Dmitri"],"committee_chairs":[],"committee_members":["Caverlee, James","Reddy, A. L. Narasimha"],"year":2012,"date_issued":"2012-02-14","date_published":"2012-02-14","updated_at":"2026-08-21T16:48:44Z","subjects":["search engines","web crawling","spam"],"languages":["en_US"],"rights":[],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"https://hdl.handle.net/1969.1/ETD-TAMU-2011-12-10235","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"source_record":{"url":"https://oaktrust.library.tamu.edu/server/oai/request?verb=GetRecord&metadataPrefix=dim&identifier=oai%3Aoaktrust.library.tamu.edu%3A1969.1%2FETD-TAMU-2011-12-10235","prefix":"dim"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Loguinov, Dmitri"]},{"key":"dc:contributor.committeemember","label":"Committee Member","values":["Caverlee, James","Reddy, A. L. Narasimha"]},{"key":"dc:creator","label":"Author","values":["Mathiharan, Siddhartha Sankaran"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2012-02-14T22:19:18Z","2012-02-16T16:19:20Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2014-01-15T07:05:32Z"]},{"key":"dc:date.issued","label":"Date","values":["2012-02-14"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Computer Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Masters"]},{"key":"thesis:degree_name","label":"Degree Name","values":["Master of Science"]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["Texas A&M University"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["search engines","web crawling","spam"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://hdl.handle.net/1969.1/ETD-TAMU-2011-12-10235"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Web crawlers encounter both finite and infinite elements during crawl. Pages and hosts can be infinitely generated using automated scripts and DNS wildcard entries. It is a challenge to rank such resources as an entire web of pages and hosts could be created to manipulate the rank of a target resource. It is crucial to be able to differentiate genuine content from spam in real-time to allocate crawl budgets. In this study, ranking algorithms to rank hosts are designed which use the finite Pay Level Domains(PLD) and IPv4 addresses. Heterogenous graphs derived from the webgraph of IRLbot are used to achieve this. PLD Supporters (PSUPP) which is the number of level-2 PLD supporters for each host on the host-host-PLD graph is the first algorithm that is studied. This is further improved by True PLD Supporters(TSUPP) which uses true egalitarian level-2 PLD supporters on the host-IP-PLD graph and DNS blacklists. It was found that support from content farms and stolen links could be eliminated by finding TSUPP. When TSUPP was applied on the host graph of IRLbot, there was less than 1% spam in the top 100,000 hosts."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Identifying Search Engine Spam Using DNS"]}]}],"canonical_facts":{"dc:contributor.advisor":["Loguinov, Dmitri"],"dc:contributor.committeemember":["Caverlee, James","Reddy, A. L. Narasimha"],"dc:creator":["Mathiharan, Siddhartha Sankaran"],"dc:date.accessioned":["2012-02-14T22:19:18Z","2012-02-16T16:19:20Z"],"dc:date.available":["2014-01-15T07:05:32Z"],"dc:date.issued":["2012-02-14"],"dc:description.abstract":["Web crawlers encounter both finite and infinite elements during crawl. Pages and hosts can be infinitely generated using automated scripts and DNS wildcard entries. It is a challenge to rank such resources as an entire web of pages and hosts could be created to manipulate the rank of a target resource. It is crucial to be able to differentiate genuine content from spam in real-time to allocate crawl budgets. In this study, ranking algorithms to rank hosts are designed which use the finite Pay Level Domains(PLD) and IPv4 addresses. Heterogenous graphs derived from the webgraph of IRLbot are used to achieve this. PLD Supporters (PSUPP) which is the number of level-2 PLD supporters for each host on the host-host-PLD graph is the first algorithm that is studied. This is further improved by True PLD Supporters(TSUPP) which uses true egalitarian level-2 PLD supporters on the host-IP-PLD graph and DNS blacklists. It was found that support from content farms and stolen links could be eliminated by finding TSUPP. When TSUPP was applied on the host graph of IRLbot, there was less than 1% spam in the top 100,000 hosts."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["https://hdl.handle.net/1969.1/ETD-TAMU-2011-12-10235"],"dc:language.iso":["en_US"],"dc:subject":["search engines","web crawling","spam"],"dc:title":["Identifying Search Engine Spam Using DNS"],"dc:type":["Thesis"],"thesis:degree_discipline":["Computer Science"],"thesis:degree_level":["Masters"],"thesis:degree_name":["Master of Science"],"thesis:institution_name":["Texas A&M University"]},"updated_at":"2026-08-21T16:48:44Z"}