# It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

> SYNTH's fully synthetic pre-training corpus collapses pre-, mid-, and post-training into one stage, achieving 10-140x token efficiency and state-of-the-art factual precision across all model sizes.

- **Source:** [arXiv](https://arxiv.org/abs/2609.37891)
- **Published:** 2026-10-10
- **Permalink:** https://picx.dev/p/wiovDP
- **Whiteboard:** https://picx.dev/p/wiovDP/image

## Summary

# It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

## Summary (Overview)

- The paper introduces **SYNTH**, the first open-source fully synthetic pre-training corpus (~80B tokens in 8 languages) derived from 58,698 Wikipedia articles, designed to collapse pre-, mid-, and post-training into a single training stage.
- The authors train a suite of models (**BAGUETTOTRON** family) ranging from 56M dense to 13B/1B-active Mixture-of-Experts, achieving competitive performance with open baselines (Gemma, Qwen, LFM) while training on **10–140× fewer tokens**.
- SYNTH-trained models achieve the **highest factual precision (FActScore) in every parameter tier**, with the 600M model reaching 79.3% precision on verifiable facts versus 65.7% for Qwen3-0.6B trained on ~36T tokens.
- The pipeline uses a **two-stage generation approach**: fine-tuned auxiliary models (query model + reasoning model) back-translate grounded seed passages into diverse task formats including memorization QA, RAG, arithmetic, creative writing, and editing.
- Domain adaptation is demonstrated in telecommunications: a 600M model fine-tuned on ~460M synthetic tokens from 3GPP standards improves TeleQnA accuracy from 41.6% to 56.7% (+15.1 points).

## Introduction and Theoretical Foundation

### Background and Motivation

Traditional LLM pre-training relies on web crawls (Common Crawl, The Pile, FineWeb), which suffer from:
- **Little content control** — inability to curate what models learn
- **Copyright concerns** — content removals and licensing issues
- **Poor alignment with downstream needs** — crawled data contains little explicit reasoning

Frontier labs have begun augmenting pre-training with synthetic data (e.g., reasoning traces from DeepSeek R1), but these datasets are proprietary and risk **model collapse** — recursive training reduces distributional support and output diversity [Shumailov et al., 2024, Dohmatob et al., 2024, Alemohammad et al., 2023].

### Theoretical Foundation: Controlled Memorization Capacity

The paper builds on controlled-environment research showing:

- **Memorization capacity**: GPT-style transformers can memorize ~3.6 bits/parameter on random bit-strings [Morris et al., 2025], with ~2 bits/parameter on synthetic biographies [Allen-Zhu and Li, 2024].
- **Data composition sensitivity**: A 1:7 useful-to-junk ratio reduces effective capacity for useful knowledge by a factor of 20 [Allen-Zhu and Li, 2024].
- **Wikipedia as seed**: A 110M model pre-trained on annotated Wikipedia matches a 10× larger baseline on entity-fact recall [Ye et al., 2026]; Llama 3.1 8B mid-trained on 1T tokens of synthetic Wikipedia-sourced material beats much larger baselines on closed-book factual QA [Lin et al., 2025b].

### Key Innovation

SYNTH replaces single-template generation (as in Phi-1.5 and Cosmopedia) with a **constraint grammar over heterogeneous task pipelines**, explicitly targeting surface-diversity collapse [Kang et al., 2025].

## Methodology

### Two-Stage Pipeline

**Stage 1 (Supervision Distillation):**
- Fine-tune two auxiliary models on data distilled from a frontier LLM:
  - **Query model**: LoRA adapter on Gemma-3-12B-base
  - **Reasoning model**: Separate auxiliary producing answers with reasoning traces
- Frozen **bge-m3 encoder** handles nearest-neighbor retrieval over seed paragraphs

**Stage 2 (Scale Generation):**
- Each seed is sampled into queries under a **constraint-prior model**
- Queries paired with retrieved neighbor paragraphs
- Routed through task-specific output adapters

### Constraint Prior Model

Constraints sampled from independent priors over six axes:

| Axis | Options |
|------|---------|
| Query type | Various question formats |
| Complexity | Difficulty levels |
| User profile | Different personas |
| Query result | 20% target refusal/correction/hedge |
| Target language | 8 languages |
| Query style | Style variations |

### Seed Corpus Composition

- **50k Wikipedia vital articles** (levels 1–5), community-curated for importance
- **8,698 specialized articles** in law, medicine, chemistry (category-tree expansion)
- **3,727 Wikibooks pages** (cooking, practical knowledge)
- **130 documents** covering model self-documentation and recent events

### Reasoning Trace Syntax

The trace uses a **stenographic syntax** (not natural-language CoT) with marker families:
- **Logical**: → (implication), ⟲ (iteration), ∴ (conclusion)
- **Epistemic**: • certain, ⃝ uncertain
- **Verification**: $\bigcirc / \bigcirc / \bigcirc$ (three-step checks)
- **Entropy markers**: ⟨H≈X.X⟩ at decision points

These dedicated tokens compress reasoning steps, letting small models emit richer traces within fixed context windows.

### Task Pipelines

1. **Memorization** (core): Query + seed paragraph + nearest-neighbor retrieval → answer with reasoning trace
2. **RAG**: Up to 10 retrieved passages with ⟨source⟩-cited targets
3. **Arithmetic**: ~3,000 Kimina templates with randomized values (using Qwen-3-8B for accuracy)
4. **Creative writing**: Lipograms, layout poems, style/persona constraints
5. **Editing**: Translation, extraction, correction, reformulation
6. **MCQ**: Closed-set questions with distractors
7. **Practical knowledge**: Wikibooks cooking recipes

### Model Suite

| Model | Parameters | Architecture | Training Tokens |
|-------|-----------|--------------|-----------------|
| MONAD | 56M | 64 layers, $d_{model}=384$ | 180B |
| BAGUETTOTRON-350M | 321M | 80 layers, $d_{model}=576$ | 200B |
| BAGUETTOTRON-600M | 594M | 48 layers, $d_{model}=1024$ | 158B |
| BAGUETTOTRON-MoE | 13.2B total / 1.05B active | 47 layers, 16 experts top-1 | 50B |

All runs use: sequence length 2048, AdamW (weight decay 0.01), 16.6% linear decay tail to 0.2% of peak LR.

## Empirical Validation / Results

### Data Quality Assessment

Using Propella-1 annotations, SYNTH occupies the **top-right quadrant** on quality/value pairings versus Nemotron-CC, FinePDFs, FineWeb-2, FineWiki, and HPLT, with the largest margins on reasoning indicators and educational value.

### Token Efficiency (Figure 4)

- **BAGUETTOTRON-600M** trails Qwen3-0.6B by only 4.7 points on multiple-choice and 0.7 on open-ended tasks, while training on **80–700× fewer tokens**
- **MONAD-56M** outperforms Gemma-3-270M and SmolLM2-360M on average multiple-choice accuracy
- Our models tie or beat Qwen on TruthfulQA, ESGenius, and FormationEval

### Data Ablations

Two identical 600M models trained on FineWiki and FinePDFs-Edu (with post-training on SmolTalk + MMLU):

- **SYNTH leads by 16–17 points** on multiple-choice and 11–14 points on open-ended tasks
- Web models stay at chance on MMLU (24.0% and 25.6%)
- **Web pre-training requires a second stage; synthetic pre-training does not**

### Reasoning Trace Ablation

Removing reasoning traces (keeping queries/answers matched):
- Multiple-choice: 41.8% vs. 42.2% (no significant change)
- Open-ended: 24.3% → 22.0% (significant drop)
- Loss concentrated in truthfulness (TruthfulQA, −9.1) and domain reasoning (NuclearQA, −10.0)

### Factual Precision (Table 1)

| Model | Tokens | S/(S+C) | macro |
|-------|--------|---------|-------|
| MONAD (56M, ours) | 180B | 42.9% | 16.3% |
| SmolLM2-360M-IT | 4T | 61.8% | 26.6% |
| LFM2.5-350M | 28T | 59.3% | 22.0% |
| **BAGUETTOTRON-350M (ours)** | **200B** | **65.1%** | **32.4%** |
| Qwen3-0.6B | ~36T | 65.7% | 31.6% |
| Phi-4-mini-instruct (3.8B) | ~5T | 77.4% | 29.5% |
| **BAGUETTOTRON-600M (ours)** | **158B** | **79.3%** | **41.7%** |
| OLMoE-1B-7B-Instruct | ~5.1T | 82.6% | 33.4% |
| DeepSeek-MoE-16B-Chat | ~2T | 82.9% | 38.3% |
| **BAGUETTOTRON-MoE (ours)** | **50B** | **81.7%** | **46.3%** |

**Key finding**: SYNTH-trained models are the most token-efficient factual learners in every parameter tier, achieving 10–140× data efficiency.

### Calibrated Abstention (Epistemic Markers)

- **BAGUETTOTRON-MoE** abstains on 67% of held-out (out-of-seed) entities vs. 20% in-seed
- OLMoE-1B-7B-Instruct abstains on only 7% of held-out entities
- Models above ~300M parameters calibrate factual precision against uncertainty markers:
  - BAGUETTOTRON-350M: 0.35 → 0.25 FActScore (confident → uncertain)
  - BAGUETTOTRON-MoE: 0.50 → 0.41
- All models commit to 15–20% fewer atomic facts on uncertain traces

### Domain Adaptation (Telecommunications)

| Metric | Base 600M | Wiki-only fine-tune | 3GPP-augmented |
|--------|-----------|---------------------|----------------|
| TeleQnA accuracy | 41.6% | — | **56.7%** (+15.1) |
| FActScore on 3GPP | 21.5% | — | **38.8%** (+17.3) |
| LLM-as-judge score | 27.8 | 58.7 | **67.8** (2.4× base) |

### Tool Calling

- BAGUETTOTRON-600M reaches **53.1% on BFCL v2**, 3.9 points above FunctionGemma-270M on ~40× fewer training tokens
- Capability comes entirely from the synthetic tool-calling data (base model emits no tool calls)

## Theoretical and Practical Implications

### Recontextualizing Scaling Laws

- **Chinchilla coefficients were fit on web text**, which is weakly aligned with small-model capabilities
- Engineered data shifts the curve: training remains productive well past the canonical 20× token-to-parameter ratio
- Confirms the Phi-series intuition that **data quality is an additional controllable axis** alongside model size and token count

### Single-Stage Training

SYNTH eliminates the need for separate instruction tuning or RLHF stages:
- Behaviors that web models only acquire in post-training (instruction following, answer formats) are written into SYNTH's pre-training corpus
- Even after post-training, web models trail by 11–17 points

### Data Releasability and Transparency

- A viable pre-training environment can be built from a small collection of 50,000 Wikipedia articles
- Open data sources (Wikipedia, Wikidata) consistently emerge as higher quality than webcrawl for seed infrastructures
- The SYNTH dataset and BAGUETTOTRON models are released under a permissive license

## Conclusion

The paper demonstrates that **fully synthetic, single-stage training** can produce competitive generalist models from a fraction of the training data, with:

1. **Superior token efficiency** (10–140× fewer tokens)
2. **Higher factual precision** with calibrated abstention
3. **Native instruction-following** without separate post-training
4. **Targeted domain adaptation** from small seed corpora

### Future Work

- **Seed coverage**: Expand beyond 58k Wikipedia articles (especially multilingual versions)
- **Cultural diversity**: Native-language seed selection beyond English-centric content
- **Capabilities**: Extend to code generation, agentic scenarios, long-horizon tasks
- **Hybrid mixes**: Combine synthetic with organic web data at scale (current work isolates synthetic contribution)

The paper opens paths for both **generalist models with significantly increased data efficiency** and **domain-specific models where no instruction or conversational data is available**.

---

_Markdown view of https://picx.dev/p/wiovDP, served by PicX — AI-generated visual whiteboard summaries of research papers._
