How Far Can Synthetic Data Take Thai OCR?

Summary (Overview)

  • Controllable document reconstruction pipeline: The authors introduce a pipeline that replaces source text in-place in existing documents while independently controlling source domain, non-text context, typeface diversity, two-dimensional layout, and handwriting glyph source.
  • Key transfer findings: Typeface diversity, two-dimensional structure, and real handwriting glyphs improve synthetic-to-real transfer; non-text page context has little consistent effect; and source-domain matching depends on training granularity (in-domain reconstruction helps page-level training but hurts crop-level training).
  • Competitive synthetic-only model: Wayu-Paxa-OCR-Zero (0.9B parameters), trained only on 45,723 synthetic pages, reduces median CER from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting relative to its base checkpoint, and outperforms Typhoon OCR v1 (7B) on all five evaluation sets.
  • Proximity to real supervision: In-domain synthetic reconstruction approaches real printed supervision under page-level training (1.82% vs. 1.31% median CER on printed pages), but a significant gap remains for handwriting and severe-error tail cases.

Introduction and Theoretical Foundation

Optical character recognition (OCR) converts document images into machine-readable text, with modern vision-language models (VLMs) performing end-to-end recognition while preserving reading order and document context. While proprietary systems (Gemini, GPT) and open models (Unlimited OCR, PaddleOCR-VL) provide strong multilingual capabilities, coverage remains uneven—particularly for less-resourced languages like Thai.

Key challenges for Thai OCR:

  • Thai's unique glyph system prevents transfer from English or Chinese models
  • PDF text extraction and OCR pseudo-labels can omit characters, reorder combining marks, and corrupt reading order
  • Manual correction is costly; open Thai datasets remain limited

Prior work gap: Recent synthetic OCR approaches (Indic pages, Manchu word images, Arabic reconstruction, historical pages) vary several generation factors simultaneously, leaving unclear whether transfer comes from layout, non-text context, fonts, or training granularity. This matters because whole-page models and detector–recognizer systems (GLM-OCR, PaddleOCR-VL) expose different amounts of document context.

Central research question: How far can synthetic data take Thai OCR?

Methodology

Document Reconstruction Pipeline

The pipeline reconstructs existing documents in-place:

  1. Source selection: Thai sources use OCR labels directly (In-Domain); non-Thai sources are translated into Thai or retain English labels (Out-of-Domain)
  2. Text erasure: Source text pixels are inpainted/erased
  3. Fit-constrained rendering: OCR labels are shaped with HarfBuzz; type size is reduced until the label fits the original region (rejecting pages that overflow at minimum size)
  4. Typeface rendering: Thai typefaces sampled from a character-weighted profile measured from 8,000 public Thai PDF pages (Table 1)
  5. Handwriting real-glyph rendering: Supported Thai characters replaced with instances from a glyph bank (~6,000 instances, 76 character classes) built from the iApp Handwriting Dataset and real Thai handwriting training splits

Experimental Design

Models: Qwen3-VL-2B-Instruct for controlled experiments; PaddleOCR-VL-1.6 (0.9B) for final model

Training settings:

  • Page-level: Model receives complete document image, predicts full page in one pass
  • Crop-level: Model recognizes individual regions from a layout detector (PP-DocLayoutV3)

Data sources:

  • Real Thai (Print): ~34,000 pages from public Thai PDFs and Common Crawl
  • Real Thai (Handwriting): ~4,000 photographed study-notebook pages
  • Out-of-Domain Synthetic: English pages from DocLayNet, Crello, PubTabNet (7.61% retain English labels)
  • In-Domain Synthetic: Reconstructed Real Thai (Print) pages

Evaluation: Character Error Rate (CER) with two aggregates—median page CER (typical page) and character-weighted mean CER (aggregate error). Three internal sets: Heldout (301 printed pages), Handwriting (200 pages), Easy Handwriting (200 legible pages). Fuzzy alignment projects predictions onto evaluation regions before scoring.

Training parameters: One epoch, AdamW, learning rate 3×10−53 \times 10^{-5}, cosine schedule, full parameter updates.

Empirical Validation / Results

Source-Property Ablation (Table 2)

Training dataRegimeHeldout Med.Handwriting Med.Easy Handwriting Med.
Out-of-Domain Syntheticpage5.0743.9938.90
– non-text contextpage4.7843.1638.07
– font diversitypage5.3449.8647.55
– two-dimensional layoutpage5.0758.4049.60
Out-of-Domain Syntheticcrop5.5249.1548.26
– non-text contextcrop4.8650.7747.81
– font diversitycrop7.0162.1361.18
– two-dimensional layoutcrop9.6070.1767.19

Key findings:

  • Removing non-text context: no consistent effect (≤1.62 points change)
  • Removing font diversity: consistent handwriting degradation (6.70–13.37 points)
  • Removing two-dimensional layout: further handwriting degradation; also hurts printed crop-level (7.01→9.60)

In-Domain vs. Out-of-Domain Reconstruction (Table 3)

Training dataRegimeHeldout Med.Handwriting Med.
Out-of-Domain Syntheticpage5.0743.99
In-Domain Syntheticpage1.8236.27
Out-of-Domain Syntheticcrop5.5249.15
In-Domain Syntheticcrop15.5952.77

Critical reversal: In-domain reconstruction helps page-level training but substantially hurts crop-level training (15.59% vs. 5.52% on Heldout). The cause remains unclear and warrants further study.

Reconstruction vs. Real Supervision (Table 4, page-level)

Training dataHeldout Med.Handwriting Med.
Qwen3-VL-2B-Instruct (baseline)14.4761.59
Out-of-Domain Synthetic5.0743.99
In-Domain Synthetic1.8236.27
Real Thai (Print)1.3136.14
Real Thai (Print + Handwriting)1.4026.05

In-domain synthetic approaches real printed supervision on typical pages (1.82% vs. 1.31% median CER) but shows a larger gap under mean CER (16.20% vs. 9.79%), indicating real supervision reduces a tail of severe errors. Real handwriting supervision remains necessary for handwriting transfer (26.05% vs. 36.27%).

Handwriting Rendering (Table 5)

Training dataRegimeHeldout Med.Handwriting Med.Easy Handwriting Med.
Out-of-Domain Syntheticpage5.0743.9938.90
+ handwriting typefacespage5.8938.6534.84
+ real glyph instancespage6.4137.9130.66
Out-of-Domain Syntheticcrop5.5249.1548.26
+ handwriting typefacescrop4.5142.4440.64
+ real glyph instancescrop4.3439.9735.66

Real glyph instances further improve handwriting medians beyond handwriting typefaces, showing real glyph variation matters. Gap to real handwriting supervision remains (26.05% on Handwriting).

Wayu-Paxa-OCR-Zero Results (Table 7)

SystemHeldout Med.Handwriting Med.Easy Handwriting Med.ThaiOCRBench Med.SEA-DocBench Med.
PaddleOCR-VL-1.6 (0.9B)6.6474.8773.7438.08.87
Wayu-Paxa-OCR-Zero (0.9B)1.2420.5514.1815.34.86
Typhoon OCR (7B)2.5443.6034.9930.69.22
Typhoon OCR 1.5 (2B)0.2119.369.026.25.81
Gemini 3.7 Flash0.0011.293.890.95.51

Wayu-Paxa-OCR-Zero outperforms Typhoon OCR (7B) on all five benchmarks despite using 0.9B parameters, nearly matches Typhoon OCR 1.5 (2B) on Handwriting, and achieves the best SEA-DocBench result among all compared systems. Training data composition: 39,534 base synthetic pages, 3,999 handwriting-focused pages, and 2,190 filled forms (45,723 total), all from public English sources with no real Thai document images.

Theoretical and Practical Implications

For synthetic data research:

  • "Realism" is not monolithic—source domain, page context, typography, spatial structure, and glyph variation have distinct and sometimes interacting effects on transfer
  • Training granularity (page vs. crop) can reverse conclusions about which source domain is preferable, highlighting the importance of matching evaluation to deployment setting
  • Typeface diversity and two-dimensional structure are more important than non-text context for out-of-distribution generalization

For Thai OCR practice:

  • Synthetic-only training can produce competitive Thai OCR without page-level OCR labels from real Thai documents, reducing the need for expensive manual annotation
  • The approach generalizes beyond internal benchmarks (ThaiOCRBench, SEA-DocBench), suggesting broad applicability
  • Handwriting remains the hardest challenge; real glyph variation helps but does not fully close the gap to real handwriting supervision

Conclusion

The paper demonstrates that synthetic reconstruction can produce competitive Thai OCR without real Thai document labels. Key determinants of transfer include typeface diversity, two-dimensional structure, and real handwriting glyphs, while non-text context has little consistent effect. The interaction between source domain and training granularity (in-domain helps page-level, hurts crop-level) remains an open question.

Future directions:

  • Understanding the cause of the source-domain reversal under crop-level training
  • Expanding the handwriting glyph bank (currently only 5,953 instances)
  • Broader coverage of document layouts, typography, and handwriting variation
  • Extending the approach to other languages with limited document annotations

Related papers