Cal Poly
Combination of a Probabilistic-Based and a Rule-Based Approach for Genealogical Record Linkage
Abstract
dc:description.abstract<p>Record linkage is the task of identifying records within one or multiple databases that refer to the same entity. Currently, there exist many different approaches for record linkage. Some approaches incorporate the use of heuristic rules, mathematical models, Markov models, or machine learning. This thesis focuses on the application of record linkage to genealogical records within family trees. Today, large collections of genealogical records are stored in databases, which may contain multiple records that refer to a single individual. Resolving duplicate genealogical records can extend our knowledge on who has lived and more complete information can be constructed by combining all information referring to an individual. Simple string matching is not a feasible option for identifying duplicate records due to inconsistencies such as typographical errors, data entry errors, and missing data.</p> <p>Record linkage algorithms can be classified under two broad categories, a rule-based or heuristic approach, or a probabilistic-based approach. The Cocktail Approach, presented by Shirley Ong Ai Pei, combines a probabilistic-based approach with a rule-based approach for record linkage. This thesis discusses a re-implementation and adoption of the Cocktail Approach to genealogical records.</p>
Degree
thesis:*- Name thesis:degree_name
- MS in Computer Science
- Discipline thesis:degree_discipline
- Computer Science
- Year dc:date.available
- 2015
Author and committee
dc:creator, dc:contributor.*- Author dc:creator
-
- Shah, Pooja P.
- Contributors dc:contributor
-
- Franz Kurfess
Subjects
dc:subject × 3Identifiers
dc:identifier.*- Identifier
- 10.15368/theses.2015.12
- OAI identifier oai:identifier
- oai:digitalcommons.calpoly.edu:theses-2464