A CNN Framework for Real-Time and Edge Deployment with High Accuracy for Lightweight Deep Learning for Bengali OCR

Accurate Bengali OCR remains challenging due to complex characters and handwriting diversity. This work presents a lightweight CNN achieving 98.29% accuracy with low computational cost.

Published in Computational Sciences

Like

Share this post

Choose a social network to share with, or copy the URL to share elsewhere

This is a representation of how your post may appear on social media. The actual post will vary between social networks

Optical Character Recognition commonly abbreviated as OCR serves as the primary technical bridge between physical handwritten documents and modern digital infrastructure (pp. 1-2). While Latin based scripts have enjoyed highly mature and commercially viable OCR solutions for decades Indic scripts present a completely different frontier of computational complexity that continues to challenge researchers globally (p. 1). Among these diverse scripts the Bengali script demands special attention due to its massive scale and cultural importance (p. 1). As the seventh most spoken language in the entire world and one of the major official languages of India Bengali is a vital medium of communication for over two hundred and thirty million people (p. 1). Effective digital processing automated data entry and seamless information retrieval in Bengali require highly accurate text digitization systems that can work reliably across various platforms (pp. 1-2). In recent years deep learning has completely revolutionized the field of computer vision replacing traditional image processing techniques that relied heavily on manual feature extraction and human designed algorithms (p. 2). Architectural paradigms like deep Convolutional Neural Networks ensemble models and hybrid CNN Transformer networks have pushed character recognition accuracy to new heights across many languages (p. 1). However these state of the art frameworks come with a severe practical drawback which is their massive computational complexity and heavy memory footprints (p. 1). They usually require large input dimensions and millions of parameters rendering them completely unsuitable for deployment on resource constrained platforms like mobile phones and embedded edge systems (p. 1). To resolve this conflict between accuracy and efficiency this study introduces a lightweight CNN architecture engineered specifically for cost effective real time handwritten Bengali character recognition (p. 1). By optimizing internal convolutional filters and shrinking the input canvas to an ultra compact forty by forty pixel resolution the proposed framework dramatically cuts down parameter sizes and floating point operations (p. 1). Despite its minimal footprint the model achieves an elite ninety eight point twenty nine percent training accuracy and a robust ninety six point forty three percent validation accuracy on a fifty class handwritten character dataset (pp. 1-2). This design proves that resource hungry hybrid networks are not a prerequisite for high precision script recognition on edge devices (p. 2). To understand why handwritten Bengali character recognition is such a complex task one must examine the intrinsic geometry and expansiveness of the script itself (p. 2). Unlike Western alphabets which consist of isolated linear characters Bengali features an intricate non linear structural topology (pp. 1-2). The script is built upon a vast alphabet consisting of eleven basic vowels and thirty nine consonants (p. 2). However the true complexity emerges from vowel modifiers known as kar consonant modifiers known as phola and compound conjunct characters known as Juktakkhor (pp. 1-2). When two or more consonants are spoken together without an intervening vowel they merge physically to form a completely new glyph (p. 1). These compound characters often bear little to no visual resemblance to their constituent parts (p. 1). A robust Bengali OCR system must therefore learn to distinguish between hundreds of highly intricate dense and visually similar graphical variations (pp. 1-2). Compounding this structural hurdle is the Shirorekha which is the distinct horizontal line that runs along the top of Bengali words binding individual characters together (pp. 1, 12). During the preprocessing and line segmentation phases of OCR the Shirorekha makes it highly difficult for automated algorithms to cleanly isolate individual characters (pp. 1, 12). Finally the inherent variability of human handwriting introduces boundless distortion across different samples (pp. 1-2). Minor variations in stroke thickness curvature pen pressure and writing angle can easily cause a conventional neural network to misclassify a character especially when processing cursive or hurried writing styles (pp. 1-2). Traditional approaches to Bengali OCR struggled with these handwriting variations because they relied on manual feature engineering where human experts had to dictate which shapes loops and edges the computer should look for (p. 2). Modern deep learning models eliminate this bottleneck by utilizing automated feature extraction allowing the neural network to autonomously discover the most optimal visual patterns during training (p. 2). Our proposed model leverages this automated capability but applies strict architectural constraints to ensure it remains lightweight enough for edge devices (pp. 1-2). The core innovation lies in the extreme reduction of the input space down to a compact forty by forty pixels (p. 1). Most conventional computer vision models utilize standard input matrices of two hundred and twenty four by two hundred and twenty four pixels (p. 12). By scaling the character inputs down to forty by forty pixels the proposed model processes over ninety six percent fewer pixels per forward pass than standard deep learning networks (pp. 1, 12). To prevent this massive reduction in spatial data from degrading the model recognition capabilities we implemented a custom Keras based CNN framework with highly optimized internal filters (pp. 2, 8). The convolutional filters are finely tuned with small specialized kernel sizes of three by three to capture the sharp curves intersections and sudden directional changes characteristic of Bengali handwriting (pp. 9, 12). The network depth and channel widths are kept deliberately narrow suppressing parameter growth and keeping the total mathematical complexity exceptionally low (pp. 1, 12). This structural efficiency allows the model to process characters at lightning speeds without draining battery power or overheating embedded hardware processors (pp. 1, 12). To rigorously validate the capabilities of this lightweight framework extensive experiments were executed on a fifty class dataset containing a total of fourteen thousand nine hundred and ninety seven handwritten Bengali character samples (pp. 1, 8). The data was cleanly partitioned to ensure a reliable evaluation including twelve thousand training images used to optimize the network weights and two thousand nine hundred and ninety seven completely unseen images used to test generalization (pp. 1, 8). During the training phase the model rapidly converged demonstrating the high compatibility of its optimized filters with the forty by forty character representations (pp. 10, 15). The final experimental results yielded extraordinary benchmarks reaching ninety eight point twenty nine percent top one training accuracy and ninety six point forty three percent validation accuracy (pp. 1-2). An architecture scoring such high marks proves that a network does not require millions of parameters or high resolution inputs to solve intricate script patterns (pp. 12, 18). The model successfully navigated the subtle geometric variances between highly similar handwritten characters demonstrating excellent resilience against the noise distortions and style irregularities present in the testing dataset, ensuring reliable system processing capabilities globally (pp. 2, 20).

https://doi.org/10.63503/j.ijssic.2026.239

Article URL: https://submissions.adroidjournals.com/index.php/ijssic/article/view/239

Please sign in or register for FREE

If you are a registered user on Research Communities by Springer Nature, please sign in