It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

Summary (Overview)

  • The paper introduces SYNTH, the first open-source fully synthetic pre-training corpus (~80B tokens in 8 languages) derived from 58,698 Wikipedia articles, designed to collapse pre-, mid-, and post-training into a single training stage.
  • The authors train a suite of models (BAGUETTOTRON family) ranging from 56M dense to 13B/1B-active Mixture-of-Experts, achieving competitive performance with open baselines (Gemma, Qwen, LFM) while training on 10–140× fewer tokens.
  • SYNTH-trained models achieve the highest factual precision (FActScore) in every parameter tier, with the 600M model reaching 79.3% precision on verifiable facts versus 65.7% for Qwen3-0.6B trained on ~36T tokens.
  • The pipeline uses a two-stage generation approach: fine-tuned auxiliary models (query model + reasoning model) back-translate grounded seed passages into diverse task formats including memorization QA, RAG, arithmetic, creative writing, and editing.
  • Domain adaptation is demonstrated in telecommunications: a 600M model fine-tuned on ~460M synthetic tokens from 3GPP standards improves TeleQnA accuracy from 41.6% to 56.7% (+15.1 points).

Introduction and Theoretical Foundation

Background and Motivation

Traditional LLM pre-training relies on web crawls (Common Crawl, The Pile, FineWeb), which suffer from:

  • Little content control — inability to curate what models learn
  • Copyright concerns — content removals and licensing issues
  • Poor alignment with downstream needs — crawled data contains little explicit reasoning

Frontier labs have begun augmenting pre-training with synthetic data (e.g., reasoning traces from DeepSeek R1), but these datasets are proprietary and risk model collapse — recursive training reduces distributional support and output diversity [Shumailov et al., 2024, Dohmatob et al., 2024, Alemohammad et al., 2023].

Theoretical Foundation: Controlled Memorization Capacity

The paper builds on controlled-environment research showing:

  • Memorization capacity: GPT-style transformers can memorize ~3.6 bits/parameter on random bit-strings [Morris et al., 2025], with ~2 bits/parameter on synthetic biographies [Allen-Zhu and Li, 2024].
  • Data composition sensitivity: A 1:7 useful-to-junk ratio reduces effective capacity for useful knowledge by a factor of 20 [Allen-Zhu and Li, 2024].
  • Wikipedia as seed: A 110M model pre-trained on annotated Wikipedia matches a 10× larger baseline on entity-fact recall [Ye et al., 2026]; Llama 3.1 8B mid-trained on 1T tokens of synthetic Wikipedia-sourced material beats much larger baselines on closed-book factual QA [Lin et al., 2025b].

Key Innovation

SYNTH replaces single-template generation (as in Phi-1.5 and Cosmopedia) with a constraint grammar over heterogeneous task pipelines, explicitly targeting surface-diversity collapse [Kang et al., 2025].

Methodology

Two-Stage Pipeline

Stage 1 (Supervision Distillation):

  • Fine-tune two auxiliary models on data distilled from a frontier LLM:
    • Query model: LoRA adapter on Gemma-3-12B-base
    • Reasoning model: Separate auxiliary producing answers with reasoning traces
  • Frozen bge-m3 encoder handles nearest-neighbor retrieval over seed paragraphs

Stage 2 (Scale Generation):

  • Each seed is sampled into queries under a constraint-prior model
  • Queries paired with retrieved neighbor paragraphs
  • Routed through task-specific output adapters

Constraint Prior Model

Constraints sampled from independent priors over six axes:

AxisOptions
Query typeVarious question formats
ComplexityDifficulty levels
User profileDifferent personas
Query result20% target refusal/correction/hedge
Target language8 languages
Query styleStyle variations

Seed Corpus Composition

  • 50k Wikipedia vital articles (levels 1–5), community-curated for importance
  • 8,698 specialized articles in law, medicine, chemistry (category-tree expansion)
  • 3,727 Wikibooks pages (cooking, practical knowledge)
  • 130 documents covering model self-documentation and recent events

Reasoning Trace Syntax

The trace uses a stenographic syntax (not natural-language CoT) with marker families:

  • Logical: → (implication), ⟲ (iteration), ∴ (conclusion)
  • Epistemic: • certain, ⃝ uncertain
  • Verification: ◯/◯/◯\bigcirc / \bigcirc / \bigcirc (three-step checks)
  • Entropy markers: ⟨H≈X.X⟩ at decision points

These dedicated tokens compress reasoning steps, letting small models emit richer traces within fixed context windows.

Task Pipelines

  1. Memorization (core): Query + seed paragraph + nearest-neighbor retrieval → answer with reasoning trace
  2. RAG: Up to 10 retrieved passages with ⟨source⟩-cited targets
  3. Arithmetic: ~3,000 Kimina templates with randomized values (using Qwen-3-8B for accuracy)
  4. Creative writing: Lipograms, layout poems, style/persona constraints
  5. Editing: Translation, extraction, correction, reformulation
  6. MCQ: Closed-set questions with distractors
  7. Practical knowledge: Wikibooks cooking recipes

Model Suite

ModelParametersArchitectureTraining Tokens
MONAD56M64 layers, dmodel=384d_{model}=384180B
BAGUETTOTRON-350M321M80 layers, dmodel=576d_{model}=576200B
BAGUETTOTRON-600M594M48 layers, dmodel=1024d_{model}=1024158B
BAGUETTOTRON-MoE13.2B total / 1.05B active47 layers, 16 experts top-150B

All runs use: sequence length 2048, AdamW (weight decay 0.01), 16.6% linear decay tail to 0.2% of peak LR.

Empirical Validation / Results

Data Quality Assessment

Using Propella-1 annotations, SYNTH occupies the top-right quadrant on quality/value pairings versus Nemotron-CC, FinePDFs, FineWeb-2, FineWiki, and HPLT, with the largest margins on reasoning indicators and educational value.

Token Efficiency (Figure 4)

  • BAGUETTOTRON-600M trails Qwen3-0.6B by only 4.7 points on multiple-choice and 0.7 on open-ended tasks, while training on 80–700× fewer tokens
  • MONAD-56M outperforms Gemma-3-270M and SmolLM2-360M on average multiple-choice accuracy
  • Our models tie or beat Qwen on TruthfulQA, ESGenius, and FormationEval

Data Ablations

Two identical 600M models trained on FineWiki and FinePDFs-Edu (with post-training on SmolTalk + MMLU):

  • SYNTH leads by 16–17 points on multiple-choice and 11–14 points on open-ended tasks
  • Web models stay at chance on MMLU (24.0% and 25.6%)
  • Web pre-training requires a second stage; synthetic pre-training does not

Reasoning Trace Ablation

Removing reasoning traces (keeping queries/answers matched):

  • Multiple-choice: 41.8% vs. 42.2% (no significant change)
  • Open-ended: 24.3% → 22.0% (significant drop)
  • Loss concentrated in truthfulness (TruthfulQA, −9.1) and domain reasoning (NuclearQA, −10.0)

Factual Precision (Table 1)

ModelTokensS/(S+C)macro
MONAD (56M, ours)180B42.9%16.3%
SmolLM2-360M-IT4T61.8%26.6%
LFM2.5-350M28T59.3%22.0%
BAGUETTOTRON-350M (ours)200B65.1%32.4%
Qwen3-0.6B~36T65.7%31.6%
Phi-4-mini-instruct (3.8B)~5T77.4%29.5%
BAGUETTOTRON-600M (ours)158B79.3%41.7%
OLMoE-1B-7B-Instruct~5.1T82.6%33.4%
DeepSeek-MoE-16B-Chat~2T82.9%38.3%
BAGUETTOTRON-MoE (ours)50B81.7%46.3%

Key finding: SYNTH-trained models are the most token-efficient factual learners in every parameter tier, achieving 10–140× data efficiency.

Calibrated Abstention (Epistemic Markers)

  • BAGUETTOTRON-MoE abstains on 67% of held-out (out-of-seed) entities vs. 20% in-seed
  • OLMoE-1B-7B-Instruct abstains on only 7% of held-out entities
  • Models above ~300M parameters calibrate factual precision against uncertainty markers:
    • BAGUETTOTRON-350M: 0.35 → 0.25 FActScore (confident → uncertain)
    • BAGUETTOTRON-MoE: 0.50 → 0.41
  • All models commit to 15–20% fewer atomic facts on uncertain traces

Domain Adaptation (Telecommunications)

MetricBase 600MWiki-only fine-tune3GPP-augmented
TeleQnA accuracy41.6%—56.7% (+15.1)
FActScore on 3GPP21.5%—38.8% (+17.3)
LLM-as-judge score27.858.767.8 (2.4× base)

Tool Calling

  • BAGUETTOTRON-600M reaches 53.1% on BFCL v2, 3.9 points above FunctionGemma-270M on ~40× fewer training tokens
  • Capability comes entirely from the synthetic tool-calling data (base model emits no tool calls)

Theoretical and Practical Implications

Recontextualizing Scaling Laws

  • Chinchilla coefficients were fit on web text, which is weakly aligned with small-model capabilities
  • Engineered data shifts the curve: training remains productive well past the canonical 20× token-to-parameter ratio
  • Confirms the Phi-series intuition that data quality is an additional controllable axis alongside model size and token count

Single-Stage Training

SYNTH eliminates the need for separate instruction tuning or RLHF stages:

  • Behaviors that web models only acquire in post-training (instruction following, answer formats) are written into SYNTH's pre-training corpus
  • Even after post-training, web models trail by 11–17 points

Data Releasability and Transparency

  • A viable pre-training environment can be built from a small collection of 50,000 Wikipedia articles
  • Open data sources (Wikipedia, Wikidata) consistently emerge as higher quality than webcrawl for seed infrastructures
  • The SYNTH dataset and BAGUETTOTRON models are released under a permissive license

Conclusion

The paper demonstrates that fully synthetic, single-stage training can produce competitive generalist models from a fraction of the training data, with:

  1. Superior token efficiency (10–140× fewer tokens)
  2. Higher factual precision with calibrated abstention
  3. Native instruction-following without separate post-training
  4. Targeted domain adaptation from small seed corpora

Future Work

  • Seed coverage: Expand beyond 58k Wikipedia articles (especially multilingual versions)
  • Cultural diversity: Native-language seed selection beyond English-centric content
  • Capabilities: Extend to code generation, agentic scenarios, long-horizon tasks
  • Hybrid mixes: Combine synthetic with organic web data at scale (current work isolates synthetic contribution)

The paper opens paths for both generalist models with significantly increased data efficiency and domain-specific models where no instruction or conversational data is available.

Related papers