{"id":{"repo_id":"mit","oai_identifier":"oai:dspace.mit.edu:1721.1/28760"},"canonical_url":"https://search.dev.ndltd.org/etd/mit/oai:dspace.mit.edu:1721.1/28760","repository":{"repo_id":"mit","name":"MIT","base_url":"https://dspace.mit.edu/oai/request"},"display":{"title":"Context identification in electronic medical records","abstract":"In order to automate data extraction from electronic medical documents, it is important to identify the correct context of the extracted information. Context in medical documents is provided by the layout of documents, which are partitioned into sections by virtue of a medical culture instilled through common practice and the training of physicians. Unfortunately, formatting and labeling is inconsistently adhered to in practice and human experts are usually required to identify sections in medical documents. A series of experiments tested the hypothesis that section identification independent of the label on sections could be achieved by using a neural network to elucidate relationships between features of sections (like size, position from start of the document) and the content characteristic of certain sections (subject-specific strings). Results showed that certain sections can be reliably identified using two different methods, and described the costs involved. The stratification of documents by document type (such as History and Physical Examination Documents or Discharge Summaries), patient diagnoses and department influenced the accuracy of identification. Future improvements suggested by the results in order to fully outline the approach were described.","abstract_html":"In order to automate data extraction from electronic medical documents, it is important to identify the correct context of the extracted information. Context in medical documents is provided by the layout of documents, which are partitioned into sections by virtue of a medical culture instilled through common practice and the training of physicians. Unfortunately, formatting and labeling is inconsistently adhered to in practice and human experts are usually required to identify sections in medical documents. A series of experiments tested the hypothesis that section identification independent of the label on sections could be achieved by using a neural network to elucidate relationships between features of sections (like size, position from start of the document) and the content characteristic of certain sections (subject-specific strings). Results showed that certain sections can be reliably identified using two different methods, and described the costs involved. The stratification of documents by document type (such as History and Physical Examination Documents or Discharge Summaries), patient diagnoses and department influenced the accuracy of identification. Future improvements suggested by the results in order to fully outline the approach were described.","abstract_has_math":false,"creators":["Stephen, Reejis, 1977-"],"institution":"Massachusetts Institute of Technology","degree_name":null,"degree_level":null,"degree_discipline":null,"degree_department":"Harvard University--MIT Division of Health Sciences and Technology.","school":null,"contributors":[],"advisors":["Aziz Boxwala."],"committee_chairs":[],"committee_members":[],"year":2004,"date_issued":"2004","date_published":"2004","updated_at":"2026-07-22T22:21:09Z","subjects":["Harvard University--MIT Division of Health Sciences and Technology."],"languages":["en_US"],"rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"rights_urls":["http://dspace.mit.edu/handle/1721.1/7582"],"identifier_entries":[]},"links":{"outbound_url":"http://hdl.handle.net/1721.1/28760","outbound_label":"Handle","outbound_source":"dc:identifier.uri"},"metadata_groups":[{"id":"people","label":"People","entries":[{"key":"dc:contributor.advisor","label":"Advisor","values":["Aziz Boxwala."]},{"key":"dc:contributor.department","label":"Department","values":["Harvard University--MIT Division of Health Sciences and Technology."]},{"key":"dc:contributor.other","label":"Dc Contributor Other","values":["Harvard University--MIT Division of Health Sciences and Technology."]},{"key":"dc:creator","label":"Author","values":["Stephen, Reejis, 1977-"]}]},{"id":"academic_context","label":"Academic Context","entries":[{"key":"dc:date.accessioned","label":"Dc Date Accessioned","values":["2005-09-27T18:11:59Z"]},{"key":"dc:date.available","label":"Dc Date Available","values":["2005-09-27T18:11:59Z"]},{"key":"dc:date.issued","label":"Date","values":["2004"]},{"key":"dc:publisher","label":"Institution","values":["Massachusetts Institute of Technology"]},{"key":"dc:type","label":"Dc Type","values":["Thesis"]}]},{"id":"subjects_keywords","label":"Subjects and Keywords","entries":[{"key":"dc:subject","label":"Dc Subject","values":["Harvard University--MIT Division of Health Sciences and Technology."]}]},{"id":"language_rights","label":"Language and Rights","entries":[{"key":"dc:language.iso","label":"Language (ISO)","values":["en_US"]},{"key":"dc:rights","label":"Dc Rights","values":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."]},{"key":"dc:rights.uri","label":"Rights URI","values":["http://dspace.mit.edu/handle/1721.1/7582"]}]},{"id":"identifiers","label":"Identifiers","entries":[{"key":"dc:identifier.uri","label":"Identifier URI","values":["http://hdl.handle.net/1721.1/28760"]}]},{"id":"additional","label":"Additional Metadata","entries":[{"key":"dc:description","label":"Description","values":["Thesis (S.M.)--Harvard-MIT Division of Health Sciences and Technology, 2004.","Includes bibliographical references (leaves 66-67)."]},{"key":"dc:description.abstract","label":"Abstract","values":["In order to automate data extraction from electronic medical documents, it is important to identify the correct context of the extracted information. Context in medical documents is provided by the layout of documents, which are partitioned into sections by virtue of a medical culture instilled through common practice and the training of physicians. Unfortunately, formatting and labeling is inconsistently adhered to in practice and human experts are usually required to identify sections in medical documents. A series of experiments tested the hypothesis that section identification independent of the label on sections could be achieved by using a neural network to elucidate relationships between features of sections (like size, position from start of the document) and the content characteristic of certain sections (subject-specific strings). Results showed that certain sections can be reliably identified using two different methods, and described the costs involved. The stratification of documents by document type (such as History and Physical Examination Documents or Discharge Summaries), patient diagnoses and department influenced the accuracy of identification. Future improvements suggested by the results in order to fully outline the approach were described."]},{"key":"dc:description.degree","label":"Dc Description Degree","values":["S.M."]},{"key":"dc:format.mimetype","label":"Dc Format Mimetype","values":["application/pdf"]},{"key":"dc:title","label":"Title","values":["Context identification in electronic medical records"]}]}],"canonical_facts":{"dc:contributor.advisor":["Aziz Boxwala."],"dc:contributor.department":["Harvard University--MIT Division of Health Sciences and Technology."],"dc:contributor.other":["Harvard University--MIT Division of Health Sciences and Technology."],"dc:creator":["Stephen, Reejis, 1977-"],"dc:date.accessioned":["2005-09-27T18:11:59Z"],"dc:date.available":["2005-09-27T18:11:59Z"],"dc:date.issued":["2004"],"dc:description":["Thesis (S.M.)--Harvard-MIT Division of Health Sciences and Technology, 2004.","Includes bibliographical references (leaves 66-67)."],"dc:description.abstract":["In order to automate data extraction from electronic medical documents, it is important to identify the correct context of the extracted information. Context in medical documents is provided by the layout of documents, which are partitioned into sections by virtue of a medical culture instilled through common practice and the training of physicians. Unfortunately, formatting and labeling is inconsistently adhered to in practice and human experts are usually required to identify sections in medical documents. A series of experiments tested the hypothesis that section identification independent of the label on sections could be achieved by using a neural network to elucidate relationships between features of sections (like size, position from start of the document) and the content characteristic of certain sections (subject-specific strings). Results showed that certain sections can be reliably identified using two different methods, and described the costs involved. The stratification of documents by document type (such as History and Physical Examination Documents or Discharge Summaries), patient diagnoses and department influenced the accuracy of identification. Future improvements suggested by the results in order to fully outline the approach were described."],"dc:description.degree":["S.M."],"dc:format.mimetype":["application/pdf"],"dc:identifier.uri":["http://hdl.handle.net/1721.1/28760"],"dc:language.iso":["en_US"],"dc:publisher":["Massachusetts Institute of Technology"],"dc:rights":["M.I.T. theses are protected by copyright. They may be viewed from this source for any purpose, but reproduction or distribution in any format is prohibited without written permission. See provided URL for inquiries about permission."],"dc:rights.uri":["http://dspace.mit.edu/handle/1721.1/7582"],"dc:subject":["Harvard University--MIT Division of Health Sciences and Technology."],"dc:title":["Context identification in electronic medical records"],"dc:type":["Thesis"]},"updated_at":"2026-07-22T22:21:09Z"}