# How Far Can Synthetic Data Take Thai OCR?

> Synthetic document reconstruction trains a competitive Thai OCR model without real Thai labels, with typeface diversity and 2D layout driving transfer more than page context.

- **Source:** [arXiv](https://arxiv.org/abs/2609.03595)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/gcvebS
- **Whiteboard:** https://picx.dev/p/gcvebS/image

## Summary

# How Far Can Synthetic Data Take Thai OCR?

## Summary (Overview)

- **Controllable document reconstruction pipeline**: The authors introduce a pipeline that replaces source text in-place in existing documents while independently controlling source domain, non-text context, typeface diversity, two-dimensional layout, and handwriting glyph source.
- **Key transfer findings**: Typeface diversity, two-dimensional structure, and real handwriting glyphs improve synthetic-to-real transfer; non-text page context has little consistent effect; and source-domain matching depends on training granularity (in-domain reconstruction helps page-level training but hurts crop-level training).
- **Competitive synthetic-only model**: Wayu-Paxa-OCR-Zero (0.9B parameters), trained only on 45,723 synthetic pages, reduces median CER from 6.64% to 1.24% on printed pages and from 74.87% to 20.55% on handwriting relative to its base checkpoint, and outperforms Typhoon OCR v1 (7B) on all five evaluation sets.
- **Proximity to real supervision**: In-domain synthetic reconstruction approaches real printed supervision under page-level training (1.82% vs. 1.31% median CER on printed pages), but a significant gap remains for handwriting and severe-error tail cases.

## Introduction and Theoretical Foundation

Optical character recognition (OCR) converts document images into machine-readable text, with modern vision-language models (VLMs) performing end-to-end recognition while preserving reading order and document context. While proprietary systems (Gemini, GPT) and open models (Unlimited OCR, PaddleOCR-VL) provide strong multilingual capabilities, coverage remains uneven—particularly for less-resourced languages like Thai.

**Key challenges for Thai OCR**:
- Thai's unique glyph system prevents transfer from English or Chinese models
- PDF text extraction and OCR pseudo-labels can omit characters, reorder combining marks, and corrupt reading order
- Manual correction is costly; open Thai datasets remain limited

**Prior work gap**: Recent synthetic OCR approaches (Indic pages, Manchu word images, Arabic reconstruction, historical pages) vary several generation factors simultaneously, leaving unclear whether transfer comes from layout, non-text context, fonts, or training granularity. This matters because whole-page models and detector–recognizer systems (GLM-OCR, PaddleOCR-VL) expose different amounts of document context.

**Central research question**: *How far can synthetic data take Thai OCR?*

## Methodology

### Document Reconstruction Pipeline

The pipeline reconstructs existing documents in-place:
1. **Source selection**: Thai sources use OCR labels directly (In-Domain); non-Thai sources are translated into Thai or retain English labels (Out-of-Domain)
2. **Text erasure**: Source text pixels are inpainted/erased
3. **Fit-constrained rendering**: OCR labels are shaped with HarfBuzz; type size is reduced until the label fits the original region (rejecting pages that overflow at minimum size)
4. **Typeface rendering**: Thai typefaces sampled from a character-weighted profile measured from 8,000 public Thai PDF pages (Table 1)
5. **Handwriting real-glyph rendering**: Supported Thai characters replaced with instances from a glyph bank (~6,000 instances, 76 character classes) built from the iApp Handwriting Dataset and real Thai handwriting training splits

### Experimental Design

**Models**: Qwen3-VL-2B-Instruct for controlled experiments; PaddleOCR-VL-1.6 (0.9B) for final model

**Training settings**:
- **Page-level**: Model receives complete document image, predicts full page in one pass
- **Crop-level**: Model recognizes individual regions from a layout detector (PP-DocLayoutV3)

**Data sources**:
- *Real Thai (Print)*: ~34,000 pages from public Thai PDFs and Common Crawl
- *Real Thai (Handwriting)*: ~4,000 photographed study-notebook pages
- *Out-of-Domain Synthetic*: English pages from DocLayNet, Crello, PubTabNet (7.61% retain English labels)
- *In-Domain Synthetic*: Reconstructed Real Thai (Print) pages

**Evaluation**: Character Error Rate (CER) with two aggregates—median page CER (typical page) and character-weighted mean CER (aggregate error). Three internal sets: Heldout (301 printed pages), Handwriting (200 pages), Easy Handwriting (200 legible pages). Fuzzy alignment projects predictions onto evaluation regions before scoring.

**Training parameters**: One epoch, AdamW, learning rate $3 \times 10^{-5}$, cosine schedule, full parameter updates.

## Empirical Validation / Results

### Source-Property Ablation (Table 2)

| Training data | Regime | Heldout Med. | Handwriting Med. | Easy Handwriting Med. |
|---|---|---|---|---|
| Out-of-Domain Synthetic | page | 5.07 | 43.99 | 38.90 |
| – non-text context | page | 4.78 | 43.16 | 38.07 |
| – font diversity | page | 5.34 | 49.86 | 47.55 |
| – two-dimensional layout | page | 5.07 | 58.40 | 49.60 |
| Out-of-Domain Synthetic | crop | 5.52 | 49.15 | 48.26 |
| – non-text context | crop | 4.86 | 50.77 | 47.81 |
| – font diversity | crop | 7.01 | 62.13 | 61.18 |
| – two-dimensional layout | crop | 9.60 | 70.17 | 67.19 |

**Key findings**:
- Removing non-text context: no consistent effect (≤1.62 points change)
- Removing font diversity: consistent handwriting degradation (6.70–13.37 points)
- Removing two-dimensional layout: further handwriting degradation; also hurts printed crop-level (7.01→9.60)

### In-Domain vs. Out-of-Domain Reconstruction (Table 3)

| Training data | Regime | Heldout Med. | Handwriting Med. |
|---|---|---|---|
| Out-of-Domain Synthetic | page | 5.07 | 43.99 |
| In-Domain Synthetic | page | **1.82** | **36.27** |
| Out-of-Domain Synthetic | crop | **5.52** | **49.15** |
| In-Domain Synthetic | crop | 15.59 | 52.77 |

**Critical reversal**: In-domain reconstruction helps page-level training but substantially hurts crop-level training (15.59% vs. 5.52% on Heldout). The cause remains unclear and warrants further study.

### Reconstruction vs. Real Supervision (Table 4, page-level)

| Training data | Heldout Med. | Handwriting Med. |
|---|---|---|
| Qwen3-VL-2B-Instruct (baseline) | 14.47 | 61.59 |
| Out-of-Domain Synthetic | 5.07 | 43.99 |
| In-Domain Synthetic | 1.82 | 36.27 |
| Real Thai (Print) | **1.31** | 36.14 |
| Real Thai (Print + Handwriting) | 1.40 | **26.05** |

In-domain synthetic approaches real printed supervision on typical pages (1.82% vs. 1.31% median CER) but shows a larger gap under mean CER (16.20% vs. 9.79%), indicating real supervision reduces a tail of severe errors. Real handwriting supervision remains necessary for handwriting transfer (26.05% vs. 36.27%).

### Handwriting Rendering (Table 5)

| Training data | Regime | Heldout Med. | Handwriting Med. | Easy Handwriting Med. |
|---|---|---|---|---|
| Out-of-Domain Synthetic | page | 5.07 | 43.99 | 38.90 |
| + handwriting typefaces | page | 5.89 | 38.65 | 34.84 |
| + real glyph instances | page | 6.41 | **37.91** | **30.66** |
| Out-of-Domain Synthetic | crop | 5.52 | 49.15 | 48.26 |
| + handwriting typefaces | crop | 4.51 | 42.44 | 40.64 |
| + real glyph instances | crop | **4.34** | **39.97** | **35.66** |

Real glyph instances further improve handwriting medians beyond handwriting typefaces, showing real glyph variation matters. Gap to real handwriting supervision remains (26.05% on Handwriting).

### Wayu-Paxa-OCR-Zero Results (Table 7)

| System | Heldout Med. | Handwriting Med. | Easy Handwriting Med. | ThaiOCRBench Med. | SEA-DocBench Med. |
|---|---|---|---|---|---|
| PaddleOCR-VL-1.6 (0.9B) | 6.64 | 74.87 | 73.74 | 38.0 | 8.87 |
| **Wayu-Paxa-OCR-Zero (0.9B)** | **1.24** | **20.55** | **14.18** | **15.3** | **4.86** |
| Typhoon OCR (7B) | 2.54 | 43.60 | 34.99 | 30.6 | 9.22 |
| Typhoon OCR 1.5 (2B) | 0.21 | 19.36 | 9.02 | 6.2 | 5.81 |
| Gemini 3.7 Flash | 0.00 | 11.29 | 3.89 | 0.9 | 5.51 |

Wayu-Paxa-OCR-Zero outperforms Typhoon OCR (7B) on all five benchmarks despite using 0.9B parameters, nearly matches Typhoon OCR 1.5 (2B) on Handwriting, and achieves the best SEA-DocBench result among all compared systems. Training data composition: 39,534 base synthetic pages, 3,999 handwriting-focused pages, and 2,190 filled forms (45,723 total), all from public English sources with no real Thai document images.

## Theoretical and Practical Implications

**For synthetic data research**:
- "Realism" is not monolithic—source domain, page context, typography, spatial structure, and glyph variation have distinct and sometimes interacting effects on transfer
- Training granularity (page vs. crop) can reverse conclusions about which source domain is preferable, highlighting the importance of matching evaluation to deployment setting
- Typeface diversity and two-dimensional structure are more important than non-text context for out-of-distribution generalization

**For Thai OCR practice**:
- Synthetic-only training can produce competitive Thai OCR without page-level OCR labels from real Thai documents, reducing the need for expensive manual annotation
- The approach generalizes beyond internal benchmarks (ThaiOCRBench, SEA-DocBench), suggesting broad applicability
- Handwriting remains the hardest challenge; real glyph variation helps but does not fully close the gap to real handwriting supervision

## Conclusion

The paper demonstrates that synthetic reconstruction can produce competitive Thai OCR without real Thai document labels. Key determinants of transfer include typeface diversity, two-dimensional structure, and real handwriting glyphs, while non-text context has little consistent effect. The interaction between source domain and training granularity (in-domain helps page-level, hurts crop-level) remains an open question.

**Future directions**:
- Understanding the cause of the source-domain reversal under crop-level training
- Expanding the handwriting glyph bank (currently only 5,953 instances)
- Broader coverage of document layouts, typography, and handwriting variation
- Extending the approach to other languages with limited document annotations

---

_Markdown view of https://picx.dev/p/gcvebS, served by PicX — AI-generated visual whiteboard summaries of research papers._
