{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/8503"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/8503","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Herodotus : a peer-to-peer Web archival system","abstract":"In this thesis, we present the design and implementation of Herodotus, a peer-to-peer web archival system. Like the Wayback Machine, a website that currently offers a web archive, Herodotus periodically crawls the world wide web and stores copies of all downloaded web content. Unlike the Wayback Machine, Herodotus does not rely on a centralized server farm. Instead, many individual nodes spread out across the Internet collaboratively perform the task of crawling and storing the content. This allows a large group of people to contribute idle computer resources to jointly achieve the goal of creating an Internet archive. Herodotus uses replication to ensure the persistence of data as nodes join and leave. Herodotus is implemented on top of Chord, a distributed peer-to-peer lookup service. It is written in C++ on FreeBSD. Our analysis based on an estimated size of the World Wide Web shows that a set of 20,000 nodes would be required to archive the entire web, assuming that each node has a typical home broadband Internet connection and contributes 100 GB of storage.","abstract_html":"In this thesis, we present the design and implementation of Herodotus, a peer-to-peer web archival system. Like the Wayback Machine, a website that currently offers a web archive, Herodotus periodically crawls the world wide web and stores copies of all downloaded web content. Unlike the Wayback Machine, Herodotus does not rely on a centralized server farm. Instead, many individual nodes spread out across the Internet collaboratively perform the task of crawling and storing the content. This allows a large group of people to contribute idle computer resources to jointly achieve the goal of creating an Internet archive. Herodotus uses replication to ensure the persistence of data as nodes join and leave. Herodotus is implemented on top of Chord, a distributed peer-to-peer lookup service. It is written in C++ on FreeBSD. Our analysis based on an estimated size of the World Wide Web shows that a set of 20,000 nodes would be required to archive the entire web, assuming that each node has a typical home broadband Internet connection and contributes 100 GB of storage.","abstract_has_math":false,"creators":["Burkard, Timo, 1979-"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science.","school":null,"contributors":[],"advisors":["Robert T. Morris."],"committee_chairs":[],"committee_members":[],"year":2002,"date_issued":"2002","date_published":"2002","updated_at":"2026-07-22T22:21:25Z","subjects":["Electrical Engineering and Computer Science."],"languages":["eng"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/8503","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Robert T. Morris."]},{"key":"dc:contributor.department","label":"Department","values":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."]},{"key":"dc:creator","label":"Author","values":["Burkard, Timo, 1979-"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2005-08-23T20:40:22Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2005-08-23T20:40:22Z"]},{"key":"dc:date.issued","label":"Date","values":["2002"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Electrical Engineering and Computer Science."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["eng"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/8503"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, June 2002.","\"May 2002.\"","Includes bibliographical references (p. 63-64)."]},{"key":"dc:description.abstract","label":"Abstract","values":["In this thesis, we present the design and implementation of Herodotus, a peer-to-peer web archival system. Like the Wayback Machine, a website that currently offers a web archive, Herodotus periodically crawls the world wide web and stores copies of all downloaded web content. Unlike the Wayback Machine, Herodotus does not rely on a centralized server farm. Instead, many individual nodes spread out across the Internet collaboratively perform the task of crawling and storing the content. This allows a large group of people to contribute idle computer resources to jointly achieve the goal of creating an Internet archive. Herodotus uses replication to ensure the persistence of data as nodes join and leave. Herodotus is implemented on top of Chord, a distributed peer-to-peer lookup service. It is written in C++ on FreeBSD. Our analysis based on an estimated size of the World Wide Web shows that a set of 20,000 nodes would be required to archive the entire web, assuming that each node has a typical home broadband Internet connection and contributes 100 GB of storage."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["M.Eng."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Herodotus : a peer-to-peer Web archival system"]}]}],"canonical_facts":{"dc:contributor.advisor":["Robert T. Morris."],"dc:contributor.department":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."],"dc:contributor.other":["Massachusetts Institute of Technology. Dept. of Electrical Engineering and Computer Science."],"dc:creator":["Burkard, Timo, 1979-"],"dc:date.accessioned":["2005-08-23T20:40:22Z"],"dc:date.available":["2005-08-23T20:40:22Z"],"dc:date.issued":["2002"],"dc:description":["Thesis (M.Eng.)--Massachusetts Institute of Technology, Dept. of Electrical Engineering and Computer Science, June 2002.","\"May 2002.\"","Includes bibliographical references (p. 63-64)."],"dc:description.abstract":["In this thesis, we present the design and implementation of Herodotus, a peer-to-peer web archival system. Like the Wayback Machine, a website that currently offers a web archive, Herodotus periodically crawls the world wide web and stores copies of all downloaded web content. Unlike the Wayback Machine, Herodotus does not rely on a centralized server farm. Instead, many individual nodes spread out across the Internet collaboratively perform the task of crawling and storing the content. This allows a large group of people to contribute idle computer resources to jointly achieve the goal of creating an Internet archive. Herodotus uses replication to ensure the persistence of data as nodes join and leave. Herodotus is implemented on top of Chord, a distributed peer-to-peer lookup service. It is written in C++ on FreeBSD. Our analysis based on an estimated size of the World Wide Web shows that a set of 20,000 nodes would be required to archive the entire web, assuming that each node has a typical home broadband Internet connection and contributes 100 GB of storage."],"dc:description.degree":["M.Eng."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["http://hdl.handle.net/1721.1/8503"],"dc:language.iso":["eng"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Electrical Engineering and Computer Science."],"dc:title":["Herodotus : a peer-to-peer Web archival system"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:25Z"}