Back close

Text Extraction from Tamil and Hindi Document Images using Open Source Optical Character Recognition Tools

Publication Type : Journal Article

Source : International Journal of Scientific Research and Technology (IJSART)

Campus : Coimbatore

School : School of Physical Sciences

Department : Mathematics

Year : 2016

Abstract : Optical Character Recognition (OCR) is a technique, which is used to extract the text from document images and converted into text format. This kind of information retrieval is called as recognition based retrieval hence that it can be edited, searched, stored more efficiently. OCR is used for many applications such as library, organization, bank cheques, number plate recognition, historical book analysis and many others applications. Various OCR tools are available for converting document images in different types of languages. The primary objective of this work is to compare the performance analysis of the three different OCR tools for extracting the text information from Tamil and Hindi document images. The OCR tools considered in this analysis are Google Docs, Free Online OCR and i2OCR. Based on the conversion accuracy it is observed that the performance of Free Online OCR is better than other OCR tools.

Cite this Research Publication : Sakila A, Vijayarani S, Text Extraction from Tamil and Hindi Document Images using Open Source Optical Character Recognition Tools, International Journal of Scientific Research and Technology (IJSART), Vol. 2, Issue 1, November 2016.

Admissions Apply Now