Advancing Arabic Text Detection: A CNN-Heuristic and EfficientNet-B1-Based Framework
DOI:
https://doi.org/10.24996/ijs.2026.67.8.29Keywords:
Text detection, Object detection, Text localization, CNN, Heuristic rules, OCR, Arabic text, EfficientNetB1Abstract
In computer vision, word detection in images is still a major challenge, especially for applications like scene interpretation, document indexing, and machine translation. Arabic script presents particular difficulties because of its cursive style, the use of diacritical marks, and the range of typefaces and orientations, which sometimes lead to errors in standard models, even though this effort has achieved great strides in Latin scripts. In this work, we propose a two-step method for Arabic Text Detection (ATD), aiming to improve accuracy without overly complex pipelines. Our approach starts with a set of heuristic rules to pre-select candidate regions. These rules rely on basic geometric and statistical cues, which helped us eliminate much of the irrelevant background noise in initial tests. To further improve the choices, we use an EfficientNet-B1 CNN. This model uses Global Average Pooling and a final Softmax layer to distinguish between text and non-text regions. The evaluation, which was conducted on a broad dataset of Arabic scene photos, revealed that combining heuristics with CNN verification enhances results significantly. For instance, we observed an increase in detection accuracy from 89.45% (heuristics alone) to 96.46% with the full pipeline. Moreover, the CNN classifier itself reached an accuracy of 97.77%. These findings confirm the robustness of our approach, particularly in cluttered visual environments.




