Modular and Interactive Preprocessing Framework for Enhanced OCR in Noisy Document Images
Loading...
Date
Journal Title
Journal ISSN
Volume Title
Publisher
Springer International Publishing AG
Series Info
Lecture Notes in Networks and Systems ; Volume 1930 LNNS , Pages 449 - 462
Scientific Journal Rankings
Orcid
Abstract
Optical Character Recognition (OCR) is the process of converting a document into searchable, retrievable, and accessible data. Poor quality in images, such as being noisy, of low-resolution, tilted, or badly lighted, significantly impairs accuracy. In the real world, sources include historic archives, receipts, handwritten notes, and mobile captures. Many of these are incomplete or incorrect due to some of the above-mentioned flaws. This work presents a modular, interactive preprocessing framework with practical image-processing steps for improving the accuracy of the optical character recognition procedure. The system will be implemented in Google Colab and will have a very friendly graphical interface through which users can upload or select images, update and change the order of the applied preprocessing filters, and instantaneously visualize the results before recognition. It shares common operations: conversion to grayscale, rescale, application of CLAHE, sharpening, median filtering, Gaussian denoising, thresholding, and contour-based text region detection. Evaluation was done on a diverse set of about 5,000 images that span different noise levels, lighting conditions, and complexities of the text. The results show that adaptive preprocessing can reduce the character error rate by up to 25.9% relative to unprocessed inputs. Compared to Tesseract and EasyOCR, improved robustness was observed across engines; scaling analyses indicated performance was scalable on both CPU and GPU set-ups. While the techniques themselves are standard, the value of this work rests in providing a practical, extensible, and accessible tool-one that bridges theoretical preprocessing research into the practical needs of OCR. This framework solves current challenges in image processing within document analysis and is a promising tool for research and education.
Description
SJR 2025
0.165
Q4
H-Index
57
Subject Area and Category:
Computer Science
Computer Networks and Communications
Signal Processing
Engineering
Control and Systems Engineering
Citation
Mohmed, Y. S., Saad, M. N., & Nassef, T. M. (2026). Modular and Interactive Preprocessing Framework for Enhanced OCR in Noisy Document Images. Lecture Notes in Networks and Systems, 449–462. https://doi.org/10.1007/978-3-032-23317-2_35
