{"id":{"repo_id":"uiuc","oai_identifier":"oai:www.ideals.illinois.edu:2142/24301"},"canonical_url":"https://search.dev.ndltd.org/etd/uiuc/oai:www.ideals.illinois.edu:2142/24301","repository":{"repo_id":"uiuc","name":"University of Illinois - Urbana-Champaign","base_url":"https://www.ideals.illinois.edu/oai-pmh"},"display":{"title":"The application of file identification, validation, and characterization tools in digital curation","abstract":"\"File format identification, characterization, and validation are considered essential processes for digital preservation and, by extension, long-term data curation. These actions are performed on data objects by humans or computers, in an attempt to identify the type of a given file, derive characterizing information that is specific to the file, and validate that the given file conforms to its type specification. The present research reviews the literature surrounding these digital preservation activities, including their theoretical basis and the publications that accompanied the formal release of tools and services designed in response to their theoretical foundation. It also reports the results from extensive tests designed to evaluate the coverage of some of the software tools developed to perform file format identification, characterization, and validation actions. Tests of these tools demonstrate that more work is needed - particularly in terms of scalable solutions - to address the expanse of digital data to be preserved and curated. The breadth of file types these tools are anticipated to handle is so great as to call into question whether a scalable solution is feasible, and, more broadly, whether such efforts will offer a meaningful return on investment. Also, these tools, which serve to provide a type of baseline reading of a file in a repository, can be easily tricked. It is possible to generate files with nothing more than a proper file extension and correct magic number and have the tools \"\"positively\"\" identify the file. This is not the same as a file that conforms to its specification, and one that could be considered valid. The ability to manipulate the results returned by these tools raises issues of identity, trust, security and risk.\"","abstract_html":"&quot;File format identification, characterization, and validation are considered essential processes for digital preservation and, by extension, long-term data curation. These actions are performed on data objects by humans or computers, in an attempt to identify the type of a given file, derive characterizing information that is specific to the file, and validate that the given file conforms to its type specification. The present research reviews the literature surrounding these digital preservation activities, including their theoretical basis and the publications that accompanied the formal release of tools and services designed in response to their theoretical foundation. It also reports the results from extensive tests designed to evaluate the coverage of some of the software tools developed to perform file format identification, characterization, and validation actions. Tests of these tools demonstrate that more work is needed - particularly in terms of scalable solutions - to address the expanse of digital data to be preserved and curated. The breadth of file types these tools are anticipated to handle is so great as to call into question whether a scalable solution is feasible, and, more broadly, whether such efforts will offer a meaningful return on investment. Also, these tools, which serve to provide a type of baseline reading of a file in a repository, can be easily tricked. It is possible to generate files with nothing more than a proper file extension and correct magic number and have the tools &quot;&quot;positively&quot;&quot; identify the file. This is not the same as a file that conforms to its specification, and one that could be considered valid. The ability to manipulate the results returned by these tools raises issues of identity, trust, security and risk.&quot;","abstract_has_math":false,"creators":["Ford, Kevin M."],"institution":"University of Illinois at Urbana-Champaign","degree_name":"M.S.","degree_level":"Thesis","degree_discipline":"Library & Information Science","degree_department":null,"school":null,"contributors":["Cragin, Melissa H.","McDonough, Jerome P."],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2011,"date_issued":"2011-05-25T14:56:50Z","date_published":"2011-05-25T14:56:50Z","updated_at":"2026-07-22T22:25:23Z","subjects":["Digital curation","Digital preservation","File identification","File validation","File characterization","Preservation tools","Preservation software"],"languages":["en"],"rights":["Copyright 2011 Kevin Ford. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. A copy of the license is available at http://creativecommons.org/licenses/by-nc-nd/3.0/"],"rights_urls":[],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/2142/24301","outbound_label":"Handle","outbound_source":"dc:identifier"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor","label":"Contributor","values":["Cragin, Melissa H.","McDonough, Jerome P."]},{"key":"dc:creator","label":"Author","values":["Ford, Kevin M."]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date","label":"Dc Date","values":["2011-05-25T14:56:50Z","2011-05"]},{"key":"thesis:degree_discipline","label":"Discipline","values":["Library & Information Science"]},{"key":"thesis:degree_level","label":"Degree Level","values":["Thesis"]},{"key":"thesis:degree_name","label":"Degree Name","values":["M.S."]},{"key":"thesis:institution_name","label":"Thesis Institution Name","values":["University of Illinois at Urbana-Champaign"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Digital curation","Digital preservation","File identification","File validation","File characterization","Preservation tools","Preservation software"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language","label":"Dc Language","values":["en"]},{"key":"dc:rights","label":"Dc Rights","values":["Copyright 2011 Kevin Ford. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. A copy of the license is available at http://creativecommons.org/licenses/by-nc-nd/3.0/"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier","label":"Identifier","values":["http://hdl.handle.net/2142/24301"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["\"File format identification, characterization, and validation are considered essential processes for digital preservation and, by extension, long-term data curation. These actions are performed on data objects by humans or computers, in an attempt to identify the type of a given file, derive characterizing information that is specific to the file, and validate that the given file conforms to its type specification. The present research reviews the literature surrounding these digital preservation activities, including their theoretical basis and the publications that accompanied the formal release of tools and services designed in response to their theoretical foundation. It also reports the results from extensive tests designed to evaluate the coverage of some of the software tools developed to perform file format identification, characterization, and validation actions. Tests of these tools demonstrate that more work is needed - particularly in terms of scalable solutions - to address the expanse of digital data to be preserved and curated. The breadth of file types these tools are anticipated to handle is so great as to call into question whether a scalable solution is feasible, and, more broadly, whether such efforts will offer a meaningful return on investment. Also, these tools, which serve to provide a type of baseline reading of a file in a repository, can be easily tricked. It is possible to generate files with nothing more than a proper file extension and correct magic number and have the tools \"\"positively\"\" identify the file. This is not the same as a file that conforms to its specification, and one that could be considered valid. The ability to manipulate the results returned by these tools raises issues of identity, trust, security and risk.\"","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2011-04-22T15:56:08Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Ford_Kevin.odt: 390908 bytes, checksum: f17a6ff57e43adb34613edef2efdfb67 (MD5) Ford_Kevin.pdf: 1005008 bytes, checksum: 0c0987b637070130ec69dfe2b1974544 (MD5)","Made available in DSpace on 2011-05-25T14:56:50Z (GMT). No. of bitstreams: 3 Ford_Kevin.pdf: 1005008 bytes, checksum: 0c0987b637070130ec69dfe2b1974544 (MD5) license.txt: 4057 bytes, checksum: 4cdca096e06cab566da27fa324f01583 (MD5) Ford_Kevin.odt: 390908 bytes, checksum: f17a6ff57e43adb34613edef2efdfb67 (MD5)"]},{"key":"dc:title","label":"Title","values":["The application of file identification, validation, and characterization tools in digital curation"]}]}],"canonical_facts":{"dc:contributor":["Cragin, Melissa H.","McDonough, Jerome P."],"dc:creator":["Ford, Kevin M."],"dc:date":["2011-05-25T14:56:50Z","2011-05"],"dc:description":["\"File format identification, characterization, and validation are considered essential processes for digital preservation and, by extension, long-term data curation. These actions are performed on data objects by humans or computers, in an attempt to identify the type of a given file, derive characterizing information that is specific to the file, and validate that the given file conforms to its type specification. The present research reviews the literature surrounding these digital preservation activities, including their theoretical basis and the publications that accompanied the formal release of tools and services designed in response to their theoretical foundation. It also reports the results from extensive tests designed to evaluate the coverage of some of the software tools developed to perform file format identification, characterization, and validation actions. Tests of these tools demonstrate that more work is needed - particularly in terms of scalable solutions - to address the expanse of digital data to be preserved and curated. The breadth of file types these tools are anticipated to handle is so great as to call into question whether a scalable solution is feasible, and, more broadly, whether such efforts will offer a meaningful return on investment. Also, these tools, which serve to provide a type of baseline reading of a file in a repository, can be easily tricked. It is possible to generate files with nothing more than a proper file extension and correct magic number and have the tools \"\"positively\"\" identify the file. This is not the same as a file that conforms to its specification, and one that could be considered valid. The ability to manipulate the results returned by these tools raises issues of identity, trust, security and risk.\"","Item withdrawn by Alexis Thompson (athmpsn1@illinois.edu) on 2011-04-22T15:56:08Z Item was in collections: University of Illinois Theses & Dissertations (ID: 1) No. of bitstreams: 2 Ford_Kevin.odt: 390908 bytes, checksum: f17a6ff57e43adb34613edef2efdfb67 (MD5) Ford_Kevin.pdf: 1005008 bytes, checksum: 0c0987b637070130ec69dfe2b1974544 (MD5)","Made available in DSpace on 2011-05-25T14:56:50Z (GMT). No. of bitstreams: 3 Ford_Kevin.pdf: 1005008 bytes, checksum: 0c0987b637070130ec69dfe2b1974544 (MD5) license.txt: 4057 bytes, checksum: 4cdca096e06cab566da27fa324f01583 (MD5) Ford_Kevin.odt: 390908 bytes, checksum: f17a6ff57e43adb34613edef2efdfb67 (MD5)"],"dc:identifier":["http://hdl.handle.net/2142/24301"],"dc:language":["en"],"dc:rights":["Copyright 2011 Kevin Ford. This work is licensed under a Creative Commons Attribution-NonCommercial-NoDerivs 3.0 Unported License. A copy of the license is available at http://creativecommons.org/licenses/by-nc-nd/3.0/"],"dc:subject":["Digital curation","Digital preservation","File identification","File validation","File characterization","Preservation tools","Preservation software"],"dc:title":["The application of file identification, validation, and characterization tools in digital curation"],"thesis:degree_discipline":["Library & Information Science"],"thesis:degree_level":["Thesis"],"thesis:degree_name":["M.S."],"thesis:institution_name":["University of Illinois at Urbana-Champaign"]},"updated_at":"2026-07-22T22:25:23Z"}