Design of a Bilingual Kannada–English OCR

Year : 2010

Abstract : India is a land of many languages and consequently one often encounters documents that contain elements in multiple languages and scripts. This chapter presents an approach towards designing a bilingual OCR that can process documents containing both English and Kannada scripts which are used by the Kannada language of the southern Indian state of Karnataka. We report an efficient script identification scheme for discriminating Kannada from Roman script. We also propose a novel segmentation and recognition scheme for Kannada, which could possibly be applied to many other Indian languages as well.

