It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs
Summary (Overview)
- The paper introduces SYNTH, the first open-source fully synthetic pre-training corpus (~80B tokens in 8 languages) derived from 58,698 Wikipedia articles, designed to collapse pre-, mid-, and post-training into a single training stage.
- The authors train a suite of models (BAGUETTOTRON family) ranging from 56M dense to 13B/1B-active Mixture-of-Experts, achieving competitive performance with open baselines (Gemma, Qwen, LFM) while training on 10–140× fewer tokens.
- SYNTH-trained models achieve the highest factual precision (FActScore) in every parameter tier, with the 600M model reaching 79.3% precision on verifiable facts versus 65.7% for Qwen3-0.6B trained on ~36T tokens.
- The pipeline uses a two-stage generation approach: fine-tuned auxiliary models (query model + reasoning model) back-translate grounded seed passages into diverse task formats including memorization QA, RAG, arithmetic, creative writing, and editing.
- Domain adaptation is demonstrated in telecommunications: a 600M model fine-tuned on ~460M synthetic tokens from 3GPP standards improves TeleQnA accuracy from 41.6% to 56.7% (+15.1 points).
Introduction and Theoretical Foundation
Background and Motivation
Traditional LLM pre-training relies on web crawls (Common Crawl, The Pile, FineWeb), which suffer from:
- Little content control — inability to curate what models learn
- Copyright concerns — content removals and licensing issues
- Poor alignment with downstream needs — crawled data contains little explicit reasoning
Frontier labs have begun augmenting pre-training with synthetic data (e.g., reasoning traces from DeepSeek R1), but these datasets are proprietary and risk model collapse — recursive training reduces distributional support and output diversity [Shumailov et al., 2024, Dohmatob et al., 2024, Alemohammad et al., 2023].
Theoretical Foundation: Controlled Memorization Capacity
The paper builds on controlled-environment research showing:
- Memorization capacity: GPT-style transformers can memorize ~3.6 bits/parameter on random bit-strings [Morris et al., 2025], with ~2 bits/parameter on synthetic biographies [Allen-Zhu and Li, 2024].
- Data composition sensitivity: A 1:7 useful-to-junk ratio reduces effective capacity for useful knowledge by a factor of 20 [Allen-Zhu and Li, 2024].
- Wikipedia as seed: A 110M model pre-trained on annotated Wikipedia matches a 10× larger baseline on entity-fact recall [Ye et al., 2026]; Llama 3.1 8B mid-trained on 1T tokens of synthetic Wikipedia-sourced material beats much larger baselines on closed-book factual QA [Lin et al., 2025b].
Key Innovation
SYNTH replaces single-template generation (as in Phi-1.5 and Cosmopedia) with a constraint grammar over heterogeneous task pipelines, explicitly targeting surface-diversity collapse [Kang et al., 2025].
Methodology
Two-Stage Pipeline
Stage 1 (Supervision Distillation):
- Fine-tune two auxiliary models on data distilled from a frontier LLM:
- Query model: LoRA adapter on Gemma-3-12B-base
- Reasoning model: Separate auxiliary producing answers with reasoning traces
- Frozen bge-m3 encoder handles nearest-neighbor retrieval over seed paragraphs
Stage 2 (Scale Generation):
- Each seed is sampled into queries under a constraint-prior model
- Queries paired with retrieved neighbor paragraphs
- Routed through task-specific output adapters
Constraint Prior Model
Constraints sampled from independent priors over six axes:
| Axis | Options |
|---|---|
| Query type | Various question formats |
| Complexity | Difficulty levels |
| User profile | Different personas |
| Query result | 20% target refusal/correction/hedge |
| Target language | 8 languages |
| Query style | Style variations |
Seed Corpus Composition
- 50k Wikipedia vital articles (levels 1–5), community-curated for importance
- 8,698 specialized articles in law, medicine, chemistry (category-tree expansion)
- 3,727 Wikibooks pages (cooking, practical knowledge)
- 130 documents covering model self-documentation and recent events
Reasoning Trace Syntax
The trace uses a stenographic syntax (not natural-language CoT) with marker families:
- Logical: → (implication), ⟲ (iteration), ∴ (conclusion)
- Epistemic: • certain, ⃝ uncertain
- Verification: (three-step checks)
- Entropy markers: ⟨H≈X.X⟩ at decision points
These dedicated tokens compress reasoning steps, letting small models emit richer traces within fixed context windows.
Task Pipelines
- Memorization (core): Query + seed paragraph + nearest-neighbor retrieval → answer with reasoning trace
- RAG: Up to 10 retrieved passages with ⟨source⟩-cited targets
- Arithmetic: ~3,000 Kimina templates with randomized values (using Qwen-3-8B for accuracy)
- Creative writing: Lipograms, layout poems, style/persona constraints
- Editing: Translation, extraction, correction, reformulation
- MCQ: Closed-set questions with distractors
- Practical knowledge: Wikibooks cooking recipes
Model Suite
| Model | Parameters | Architecture | Training Tokens |
|---|---|---|---|
| MONAD | 56M | 64 layers, | 180B |
| BAGUETTOTRON-350M | 321M | 80 layers, | 200B |
| BAGUETTOTRON-600M | 594M | 48 layers, | 158B |
| BAGUETTOTRON-MoE | 13.2B total / 1.05B active | 47 layers, 16 experts top-1 | 50B |
All runs use: sequence length 2048, AdamW (weight decay 0.01), 16.6% linear decay tail to 0.2% of peak LR.
Empirical Validation / Results
Data Quality Assessment
Using Propella-1 annotations, SYNTH occupies the top-right quadrant on quality/value pairings versus Nemotron-CC, FinePDFs, FineWeb-2, FineWiki, and HPLT, with the largest margins on reasoning indicators and educational value.
Token Efficiency (Figure 4)
- BAGUETTOTRON-600M trails Qwen3-0.6B by only 4.7 points on multiple-choice and 0.7 on open-ended tasks, while training on 80–700× fewer tokens
- MONAD-56M outperforms Gemma-3-270M and SmolLM2-360M on average multiple-choice accuracy
- Our models tie or beat Qwen on TruthfulQA, ESGenius, and FormationEval
Data Ablations
Two identical 600M models trained on FineWiki and FinePDFs-Edu (with post-training on SmolTalk + MMLU):
- SYNTH leads by 16–17 points on multiple-choice and 11–14 points on open-ended tasks
- Web models stay at chance on MMLU (24.0% and 25.6%)
- Web pre-training requires a second stage; synthetic pre-training does not
Reasoning Trace Ablation
Removing reasoning traces (keeping queries/answers matched):
- Multiple-choice: 41.8% vs. 42.2% (no significant change)
- Open-ended: 24.3% → 22.0% (significant drop)
- Loss concentrated in truthfulness (TruthfulQA, −9.1) and domain reasoning (NuclearQA, −10.0)
Factual Precision (Table 1)
| Model | Tokens | S/(S+C) | macro |
|---|---|---|---|
| MONAD (56M, ours) | 180B | 42.9% | 16.3% |
| SmolLM2-360M-IT | 4T | 61.8% | 26.6% |
| LFM2.5-350M | 28T | 59.3% | 22.0% |
| BAGUETTOTRON-350M (ours) | 200B | 65.1% | 32.4% |
| Qwen3-0.6B | ~36T | 65.7% | 31.6% |
| Phi-4-mini-instruct (3.8B) | ~5T | 77.4% | 29.5% |
| BAGUETTOTRON-600M (ours) | 158B | 79.3% | 41.7% |
| OLMoE-1B-7B-Instruct | ~5.1T | 82.6% | 33.4% |
| DeepSeek-MoE-16B-Chat | ~2T | 82.9% | 38.3% |
| BAGUETTOTRON-MoE (ours) | 50B | 81.7% | 46.3% |
Key finding: SYNTH-trained models are the most token-efficient factual learners in every parameter tier, achieving 10–140× data efficiency.
Calibrated Abstention (Epistemic Markers)
- BAGUETTOTRON-MoE abstains on 67% of held-out (out-of-seed) entities vs. 20% in-seed
- OLMoE-1B-7B-Instruct abstains on only 7% of held-out entities
- Models above ~300M parameters calibrate factual precision against uncertainty markers:
- BAGUETTOTRON-350M: 0.35 → 0.25 FActScore (confident → uncertain)
- BAGUETTOTRON-MoE: 0.50 → 0.41
- All models commit to 15–20% fewer atomic facts on uncertain traces
Domain Adaptation (Telecommunications)
| Metric | Base 600M | Wiki-only fine-tune | 3GPP-augmented |
|---|---|---|---|
| TeleQnA accuracy | 41.6% | — | 56.7% (+15.1) |
| FActScore on 3GPP | 21.5% | — | 38.8% (+17.3) |
| LLM-as-judge score | 27.8 | 58.7 | 67.8 (2.4× base) |
Tool Calling
- BAGUETTOTRON-600M reaches 53.1% on BFCL v2, 3.9 points above FunctionGemma-270M on ~40× fewer training tokens
- Capability comes entirely from the synthetic tool-calling data (base model emits no tool calls)
Theoretical and Practical Implications
Recontextualizing Scaling Laws
- Chinchilla coefficients were fit on web text, which is weakly aligned with small-model capabilities
- Engineered data shifts the curve: training remains productive well past the canonical 20× token-to-parameter ratio
- Confirms the Phi-series intuition that data quality is an additional controllable axis alongside model size and token count
Single-Stage Training
SYNTH eliminates the need for separate instruction tuning or RLHF stages:
- Behaviors that web models only acquire in post-training (instruction following, answer formats) are written into SYNTH's pre-training corpus
- Even after post-training, web models trail by 11–17 points
Data Releasability and Transparency
- A viable pre-training environment can be built from a small collection of 50,000 Wikipedia articles
- Open data sources (Wikipedia, Wikidata) consistently emerge as higher quality than webcrawl for seed infrastructures
- The SYNTH dataset and BAGUETTOTRON models are released under a permissive license
Conclusion
The paper demonstrates that fully synthetic, single-stage training can produce competitive generalist models from a fraction of the training data, with:
- Superior token efficiency (10–140× fewer tokens)
- Higher factual precision with calibrated abstention
- Native instruction-following without separate post-training
- Targeted domain adaptation from small seed corpora
Future Work
- Seed coverage: Expand beyond 58k Wikipedia articles (especially multilingual versions)
- Cultural diversity: Native-language seed selection beyond English-centric content
- Capabilities: Extend to code generation, agentic scenarios, long-horizon tasks
- Hybrid mixes: Combine synthetic with organic web data at scale (current work isolates synthetic contribution)
The paper opens paths for both generalist models with significantly increased data efficiency and domain-specific models where no instruction or conversational data is available.
Related papers
- hacktrace: behavior-supervised detection of reward hacking during code generation
HACKTRACE detects reward hacking in coding agents from activations already computed during generation, cutting cheating from 85% to under 5% with only 8 ms overhead.
- CATCH: A Controllable Analysis Testbed for Reward Hacking in Coding RL
CATCH is a controllable coding-RL testbed revealing that chain-of-thought monitors suppress reward hacking initially but erode as policies learn to mislead them with code comments.
- One Skill Too Many: How Co-Installed Skills Conflict in Coding Agents
Co-installed coding-agent skills that do the same job reduce the installed skill's usage by 19.9 percentage points without lowering task completion, a conflict decided at the first skill read.