A Study of Style Effects on OCR Errors in the MEDLINE Database

Penny Garrison, Diane Davis, Tim Andersen, Elisa Barney Smith

Research output: Contribution to journalArticlepeer-review

Abstract

The National Library of Medicine has developed a system for the automatic extraction of data from scanned journal articles to populate the MEDLINE database. Although the 5-engine OCR system used in this process exhibits good performance overall, it does make errors in character recognition that must be corrected in order for the process to achieve the requisite accuracy. The correction process works by feeding words that have characters with less than 100% confidence (as determined automatically by the OCR engine) to a human operator who then must manually verify the word or correct the error. The majority of these errors are contained in the affiliation information zone where the characters are in italics or small fonts. Therefore only affiliation information data is used in this research. This paper examines the correlation between OCR errors and various character attributes in the MEDLINE database, such as font size, italics, bold, etc. and OCR confidence levels. The motivation for this research is that if a correlation between the character style and types of errors exists it should be possible to use this information to improve operator productivity by increasing the probability that the correct word option is presented to the human editor. We have determined that this correlation exists, in particular for the case of characters with diacritics.

Original languageAmerican English
Article number04
Pages (from-to)28-36
Number of pages9
JournalElectrical and Computer Engineering Faculty Publications and Presentations
Volume5676
DOIs
StatePublished - 19 Jan 2005
EventProceedings of SPIE-IS and T Electronic Imaging - Document Recognition and Retrieval XII - San Jose, CA, United States
Duration: 19 Jan 200520 Jan 2005

Keywords

  • OCR
  • OCR confidence
  • error correction
  • style effects

EGS Disciplines

  • Electrical and Computer Engineering

Fingerprint

Dive into the research topics of 'A Study of Style Effects on OCR Errors in the MEDLINE Database'. Together they form a unique fingerprint.

Cite this