Preprint
Large Language Models

Persian Pixel: A large-scale synthetic OCR dataset for Persian language

Pouria Mahdi, Haq Nawaz Malik
July 22, 2026

0

Citations

0

Influential Citations

Venue

2026

Year

Abstract

Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries. This gap arises from two fundamental challenges: the intrinsic complexity of the Perso-Arabic writing system and the limited availability of large-scale, high-quality annotated datasets. Persian script exhibits obligatory cursive connectivity, context-dependent glyph shaping, extensive ligatures, diacritic placement, and stylistic variation across writing forms such as Naskh and Nastaliq, all of which significantly complicate text recognition. At the same time, the high cost and labor-intensive nature of manual annotation have created a persistent data bottleneck, limiting the development of robust OCR systems and slowing progress in Persian document digitization.In this paper, we introduce Persian Pixel, a comprehensive synthetic OCR dataset specifically designed to address these challenges. Comprising over 343,000 high-fidelity image text pairs, the dataset spans sentence, paragraph, and full-page document layouts generated from a carefully curated seven-million-word Persian corpus using the SynthOCR-Gen rendering framework. The generation pipeline faithfully models the typographic characteristics of Persian script, including contextual character joining, positional glyph variants, diacritic placement, and multiple representative Persian typefaces. To bridge the synthetic-to-real domain gap, the rendered images are further enriched with more than twenty-five stochastic degradation models that emulate realistic document acquisition artifacts, including ink bleed, paper aging, blur, illumination variation, scanner imperfections, compression artifacts, and multiple noise processes.By overcoming the long-standing scarcity of annotated Persian OCR data, Persian Pixel provides a scalable and openly available resource for training and fine-tuning modern OCR architectures, including transformer-based models such as TrOCR and Donut. The dataset establishes a strong foundation for research in Persian document analysis, historical manuscript digitization, and end-to-end document understanding, while demonstrating that programmatic synthetic data generation offers a practical, cost-effective, and scalable alternative to manual annotation for advancing OCR in low-resource and typographically complex scripts.

Analysis

Why This Paper Matters

Persian, spoken by over 110 million people, remains underserved in optical character recognition (OCR) compared to Latin-script languages. The Perso-Arabic script's inherent complexity—cursive connectivity, context-dependent glyph shapes, extensive ligatures, and diacritic placement—poses significant challenges for text recognition. Moreover, the high cost of manual annotation has created a persistent data bottleneck, hindering the development of robust OCR systems. This paper introduces Persian Pixel, a large-scale synthetic dataset that directly tackles these issues by providing over 343,000 high-fidelity image-text pairs. By offering a scalable and openly available resource, it lowers the barrier to entry for researchers and practitioners working on Persian document digitization, a critical step for preserving cultural heritage and enabling digital services in Persian-speaking regions.

The work is particularly timely given the rise of transformer-based OCR architectures (e.g., TrOCR, Donut) that require large amounts of annotated data. Persian Pixel fills this gap with programmatically generated data, demonstrating that synthetic data can be a practical alternative to manual annotation for typographically complex scripts. This approach has implications beyond Persian, offering a template for other low-resource languages with intricate writing systems.

Technical Contributions

  • Large-scale synthetic dataset: Persian Pixel comprises over 343,000 image-text pairs spanning sentences, paragraphs, and full-page layouts, generated from a curated seven-million-word Persian corpus.
  • SynthOCR-Gen rendering framework: This pipeline faithfully models Persian script typography, including contextual character joining, positional glyph variants, diacritic placement, and multiple representative typefaces (e.g., Naskh, Nastaliq).
  • Domain gap mitigation: More than 25 stochastic degradation models simulate realistic document artifacts such as ink bleed, paper aging, blur, illumination variation, scanner imperfections, compression artifacts, and noise, helping bridge the synthetic-to-real gap.
  • Compatibility with modern architectures: The dataset is designed for training and fine-tuning transformer-based OCR models like TrOCR and Donut, enabling end-to-end document understanding.

Results

The abstract does not report quantitative OCR performance metrics (e.g., character error rate, word accuracy) or comparisons with existing datasets or models. The primary result is the creation of the dataset itself—over 343,000 image-text pairs—and the demonstration that synthetic data generation can produce a large-scale resource for Persian OCR. Future work would need to evaluate the dataset's effectiveness by training models and measuring recognition accuracy on real-world Persian documents.

Significance

Persian Pixel addresses a critical data scarcity issue for Persian OCR, a field that has lagged behind due to script complexity and annotation costs. By providing a scalable, openly available synthetic dataset, it enables research in Persian document analysis, historical manuscript digitization, and end-to-end document understanding. The methodology—programmatic synthetic data generation with realistic degradation—offers a cost-effective and scalable alternative to manual annotation, with potential applicability to other low-resource and typographically complex scripts. This work could accelerate progress in digitizing Persian cultural heritage and improving accessibility of Persian-language documents.