Modular and Interactive Preprocessing Framework for Enhanced OCR in Noisy Document Images

dc.AffiliationOctober University for modern sciences and Arts MSA
dc.contributor.authorYehia Samir Mohmed
dc.contributor.authorMohamed N. Saad
dc.contributor.authorTamer M. Nassef
dc.date.accessioned2026-09-05T09:14:46Z
dc.date.issued2026-07-02
dc.descriptionSJR 2025 0.165 Q4 H-Index 57 Subject Area and Category: Computer Science Computer Networks and Communications Signal Processing Engineering Control and Systems Engineering
dc.description.abstractOptical Character Recognition (OCR) is the process of converting a document into searchable, retrievable, and accessible data. Poor quality in images, such as being noisy, of low-resolution, tilted, or badly lighted, significantly impairs accuracy. In the real world, sources include historic archives, receipts, handwritten notes, and mobile captures. Many of these are incomplete or incorrect due to some of the above-mentioned flaws. This work presents a modular, interactive preprocessing framework with practical image-processing steps for improving the accuracy of the optical character recognition procedure. The system will be implemented in Google Colab and will have a very friendly graphical interface through which users can upload or select images, update and change the order of the applied preprocessing filters, and instantaneously visualize the results before recognition. It shares common operations: conversion to grayscale, rescale, application of CLAHE, sharpening, median filtering, Gaussian denoising, thresholding, and contour-based text region detection. Evaluation was done on a diverse set of about 5,000 images that span different noise levels, lighting conditions, and complexities of the text. The results show that adaptive preprocessing can reduce the character error rate by up to 25.9% relative to unprocessed inputs. Compared to Tesseract and EasyOCR, improved robustness was observed across engines; scaling analyses indicated performance was scalable on both CPU and GPU set-ups. While the techniques themselves are standard, the value of this work rests in providing a practical, extensible, and accessible tool-one that bridges theoretical preprocessing research into the practical needs of OCR. This framework solves current challenges in image processing within document analysis and is a promising tool for research and education.
dc.description.urihttps://www.scimagojr.com/journalsearch.php?q=21100901469&tip=sid&clean=0
dc.identifier.citationMohmed, Y. S., Saad, M. N., & Nassef, T. M. (2026). Modular and Interactive Preprocessing Framework for Enhanced OCR in Noisy Document Images. Lecture Notes in Networks and Systems, 449–462. https://doi.org/10.1007/978-3-032-23317-2_35
dc.identifier.doihttps://doi.org/10.1007/978-3-032-23317-2_35
dc.identifier.otherhttps://doi.org/10.1007/978-3-032-23317-2_35
dc.identifier.urihttps://repository.msa.edu.eg/handle/123456789/6840
dc.language.isoen_US
dc.publisherSpringer International Publishing AG
dc.relation.ispartofseriesLecture Notes in Networks and Systems ; Volume 1930 LNNS , Pages 449 - 462
dc.subjectadaptive filtering
dc.subjectapplied image processing
dc.subjectcharacter error rate (CER)
dc.subjectdocument digitization
dc.subjectimage preprocessing
dc.subjectnoisy document images
dc.subjectOptical character recognition (OCR)
dc.titleModular and Interactive Preprocessing Framework for Enhanced OCR in Noisy Document Images
dc.typeArticle

Files

Original bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
IMG-20231214-WA0000.jpg
Size:
16.8 KB
Format:
Joint Photographic Experts Group/JPEG File Interchange Format (JFIF)

License bundle

Now showing 1 - 1 of 1
Loading...
Thumbnail Image
Name:
license.txt
Size:
51 B
Format:
Item-specific license agreed upon to submission
Description: