{"id":{"repo_id":"de-montfort","oai_identifier":"oai:dora.dmu.ac.uk:2086/24509"},"canonical_url":"https://search.dev.ndltd.org/etd/de-montfort/oai:dora.dmu.ac.uk:2086/24509","repository":{"repo_id":"de-montfort","name":"De Montfort University","base_url":"https://dora.dmu.ac.uk/server/oai/request"},"display":{"title":"Automated identification of press variants in old documents","abstract":"Collation is the comparison of textual content to identify variations within the texts being compared. This involves comparing the texts word by word and character by character. Throughout history, collation has been used in a variety of application areas, using the naked eye, mechanical machines, as well as advanced automated collation methods that rely on software tools. The main objective of this research is to develop a fully automated system that can detect textual variations between copies of the same book. The system includes five main steps: pre-processing, segmentation, post-processing, feature extraction, and classification using a K-NN classifier. The post-processing step, includes a new technique to solve the character co-existence problem. It consists of counting the number of black pixels in all detected objects in the character image and eliminating objects with a small number of black pixels. This technique achieves an accuracy rate of eliminating unwanted objects from the character image of more than 90%. Another problem addressed in this research is detecting extra lines and extra words that may appear in the texts being compared. To solve this issue, a new technique was provided that uses the distance between the first and last black pixel to determine if there is an additional word. The testing was done using \"The Tragedy of Hamlet\" by William Shakespeare. The results showed that the integration of a K-NN classifier with feature extraction algorithms (Zoning, Template Matching, Crossings, Theta Distribution, Projection Profile) led to higher accuracy scores of character matching (classifying each character to the correct class) at 84% compared to the Calamari OCR system at 72%, and also a higher accuracy of textual variants detection at 88% compared to the Calamari OCR system at 73%.","abstract_html":"Collation is the comparison of textual content to identify variations within the texts being compared. This involves comparing the texts word by word and character by character. Throughout history, collation has been used in a variety of application areas, using the naked eye, mechanical machines, as well as advanced automated collation methods that rely on software tools. The main objective of this research is to develop a fully automated system that can detect textual variations between copies of the same book. The system includes five main steps: pre-processing, segmentation, post-processing, feature extraction, and classification using a K-NN classifier. The post-processing step, includes a new technique to solve the character co-existence problem. It consists of counting the number of black pixels in all detected objects in the character image and eliminating objects with a small number of black pixels. This technique achieves an accuracy rate of eliminating unwanted objects from the character image of more than 90%. Another problem addressed in this research is detecting extra lines and extra words that may appear in the texts being compared. To solve this issue, a new technique was provided that uses the distance between the first and last black pixel to determine if there is an additional word. The testing was done using &quot;The Tragedy of Hamlet&quot; by William Shakespeare. The results showed that the integration of a K-NN classifier with feature extraction algorithms (Zoning, Template Matching, Crossings, Theta Distribution, Projection Profile) led to higher accuracy scores of character matching (classifying each character to the correct class) at 84% compared to the Calamari OCR system at 72%, and also a higher accuracy of textual variants detection at 88% compared to the Calamari OCR system at 73%.","abstract_has_math":false,"creators":["Al-Ibaisi, Mohammad"],"institution":"De Montfort University","degree_name":"PhD","degree_level":"Doctoral","degree_discipline":null,"degree_department":null,"school":null,"contributors":[],"advisors":[],"committee_chairs":[],"committee_members":[],"year":2023,"date_issued":"2023-01","date_published":"2023-01","updated_at":"2026-07-24T06:18:37Z","subjects":[],"languages":[],"rights":[],"rights_urls":["https://dora.dmu.ac.uk/bitstreams/068158df-ecd8-46b9-a9fb-4720298c91c1/download"],"identifier_entries":[]},"links":{"outbound_url":null,"outbound_label":null,"outbound_source":null},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:creator","label":"Author","values":["Al-Ibaisi, Mohammad"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.issued","label":"Date","values":["2023-01"]},{"key":"dc:publisher.department","label":"Dc Publisher Department","values":["Faculty of Computing, Engineering and Media"]},{"key":"dc:publisher.institution","label":"Dc Publisher Institution","values":["De Montfort University"]},{"key":"dc:relation.isreferencedby","label":"Dc Relation Isreferencedby","values":["https://hdl.handle.net/2086/24509"]},{"key":"dc:type","label":"Dc Type","values":["Thesis or dissertation"]},{"key":"dc:type.qualificationlevel","label":"Dc Type Qualificationlevel","values":["Doctoral"]},{"key":"dc:type.qualificationname","label":"Dc Type Qualificationname","values":["PhD"]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:rights","label":"Dc Rights","values":["https://dora.dmu.ac.uk/bitstreams/068158df-ecd8-46b9-a9fb-4720298c91c1/download"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["https://dora.dmu.ac.uk/bitstreams/277b1e30-f8e4-43e8-a1c5-6897478302e5/download"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description.abstract","label":"Abstract","values":["Collation is the comparison of textual content to identify variations within the texts being compared. This involves comparing the texts word by word and character by character. Throughout history, collation has been used in a variety of application areas, using the naked eye, mechanical machines, as well as advanced automated collation methods that rely on software tools. The main objective of this research is to develop a fully automated system that can detect textual variations between copies of the same book. The system includes five main steps: pre-processing, segmentation, post-processing, feature extraction, and classification using a K-NN classifier. The post-processing step, includes a new technique to solve the character co-existence problem. It consists of counting the number of black pixels in all detected objects in the character image and eliminating objects with a small number of black pixels. This technique achieves an accuracy rate of eliminating unwanted objects from the character image of more than 90%. Another problem addressed in this research is detecting extra lines and extra words that may appear in the texts being compared. To solve this issue, a new technique was provided that uses the distance between the first and last black pixel to determine if there is an additional word. The testing was done using \"The Tragedy of Hamlet\" by William Shakespeare. The results showed that the integration of a K-NN classifier with feature extraction algorithms (Zoning, Template Matching, Crossings, Theta Distribution, Projection Profile) led to higher accuracy scores of character matching (classifying each character to the correct class) at 84% compared to the Calamari OCR system at 72%, and also a higher accuracy of textual variants detection at 88% compared to the Calamari OCR system at 73%."]},{"key":"dc:format.checksum.md5","label":"Dc Format Checksum Md5","values":["fbe5f0311b008c1a0de3fe4fd5bf4a6a","bd41181d9a4c38b5ebacc69a027024d9","ac3eaebe4b3c34fd827b8f66adf24ca6"]},{"key":"dc:title","label":"Title","values":["Automated identification of press variants in old documents"]}]}],"canonical_facts":{"dc:creator":["Al-Ibaisi, Mohammad"],"dc:date.issued":["2023-01"],"dc:description.abstract":["Collation is the comparison of textual content to identify variations within the texts being compared. This involves comparing the texts word by word and character by character. Throughout history, collation has been used in a variety of application areas, using the naked eye, mechanical machines, as well as advanced automated collation methods that rely on software tools. The main objective of this research is to develop a fully automated system that can detect textual variations between copies of the same book. The system includes five main steps: pre-processing, segmentation, post-processing, feature extraction, and classification using a K-NN classifier. The post-processing step, includes a new technique to solve the character co-existence problem. It consists of counting the number of black pixels in all detected objects in the character image and eliminating objects with a small number of black pixels. This technique achieves an accuracy rate of eliminating unwanted objects from the character image of more than 90%. Another problem addressed in this research is detecting extra lines and extra words that may appear in the texts being compared. To solve this issue, a new technique was provided that uses the distance between the first and last black pixel to determine if there is an additional word. The testing was done using \"The Tragedy of Hamlet\" by William Shakespeare. The results showed that the integration of a K-NN classifier with feature extraction algorithms (Zoning, Template Matching, Crossings, Theta Distribution, Projection Profile) led to higher accuracy scores of character matching (classifying each character to the correct class) at 84% compared to the Calamari OCR system at 72%, and also a higher accuracy of textual variants detection at 88% compared to the Calamari OCR system at 73%."],"dc:format.checksum.md5":["fbe5f0311b008c1a0de3fe4fd5bf4a6a","bd41181d9a4c38b5ebacc69a027024d9","ac3eaebe4b3c34fd827b8f66adf24ca6"],"dc:identifier.uri":["https://dora.dmu.ac.uk/bitstreams/277b1e30-f8e4-43e8-a1c5-6897478302e5/download"],"dc:publisher.department":["Faculty of Computing, Engineering and Media"],"dc:publisher.institution":["De Montfort University"],"dc:relation.isreferencedby":["https://hdl.handle.net/2086/24509"],"dc:rights":["https://dora.dmu.ac.uk/bitstreams/068158df-ecd8-46b9-a9fb-4720298c91c1/download"],"dc:title":["Automated identification of press variants in old documents"],"dc:type":["Thesis or dissertation"],"dc:type.qualificationlevel":["Doctoral"],"dc:type.qualificationname":["PhD"]},"updated_at":"2026-07-24T06:18:37Z"}