NanoOCR

An ultra-lightweight, high-speed Optical Character Recognition (OCR) engine built from scratch for edge devices and embedded systems.

pip install nano-ocr
GitHub Repository View on PyPI

18.3 MB

Model Size

4.8M

Parameters

1.74 ms

T4 GPU Latency

575

Frames Per Second

Architecture & Engineering

NanoOCR is built on a CRNN (Convolutional Recurrent Neural Network) architecture combined with CTC (Connectionist Temporal Classification) decoding.

  • Feature Extraction: A 7-layer VGG-style CNN utilizing asymmetric max-pooling to capture tall, narrow character features efficiently.
  • Sequence Modeling: A Bidirectional LSTM (hidden size 256) to understand the sequential context of characters in a word.
  • Loss & Decoding: CTC Loss allows the network to predict text directly from unsegmented images.

The Memory-Efficient Streaming Pipeline

Training an OCR model on millions of images typically crashes cloud hardware due to RAM and disk limits. To bypass Google Colab's hardware constraints, NanoOCR was trained using a custom iterable streaming pipeline. Instead of downloading datasets locally, the pipeline streams data chunk-by-chunk (like Netflix), keeping the memory footprint near zero while continuously feeding the GPU.

The 3-Phase Training Journey

Phase 1: Proof of Concept

Trained on 320,000 images. Proved the architecture could successfully learn character alignments, achieving an initial 56.27% accuracy on unseen IIIT5k data.

Phase 2: Foundation Model

Scaled up to 2.32 Million images (MJSynth dataset) using the streaming pipeline. Accuracy jumped to 71.57% on IIIT5k, mastering broad typography and synthetic fonts.

Phase 3: Real-World Fine-Tuning

Utilized Layer-Freezing. By freezing the core CNN layers and only fine-tuning the BiLSTM/Head on 2,000 real-world scene text images (IIIT5k), the model adapted to messy real-world data without catastrophic forgetting of its 2M image foundation.

Comprehensive Evaluation

Dataset Accuracy

  • ICDAR 2013 81.28%
  • MJSynth (Unseen) 79.20%
  • IIIT5k 78.37%
  • ICDAR 2015 (Heavy noise) 46.41%

Error Metrics

  • CTC Loss (ICDAR '13) 0.44
  • CER (ICDAR '13) 6.73%
  • CER (IIIT5k) 8.56%

*CER (Character Error Rate) indicates that even when an entire word is predicted "wrong", the model typically only misses a single character.

Error Typology (15,000 Character Sample)

602 Substitutions (e.g. '0' vs 'O') 628 Deletions (Missing thin letters) 77 Insertions (Hallucinations)

Confidence Calibration: The engine is self-aware. It averages 95.7% confidence on correct predictions, and drops to 76.7% when incorrect, allowing developers to build safe failure thresholds.

Applications & Edge Computing

At just 18.3 MB and requiring zero heavy dependencies, NanoOCR is purpose-built for environments where cloud OCR is impossible or too expensive:

Raspberry Pi & IoT

Runs natively on ARM devices for smart cameras and meters.

Mobile Offline Apps

Small enough to bundle inside iOS/Android applications.

Robotics Pipeline

1.74ms latency ensures text extraction won't bottleneck live video feeds.

Help Make NanoOCR More Efficient

NanoOCR is fully open-source. Whether it is improving the dataset, shrinking the weights via quantization, or optimizing the PyTorch pipeline, contributions are welcome!

Submit a Pull Request