Full text not available for this paper

Summary (Overview)

  • ReScraper is a unified 0.6B language model that replaces the entire heuristic pipeline of HTML scraping and rule-based cleaning for LLM pretraining data curation.
  • It performs four operations in a single pass: extract (main content), keep, edit (remove noisy lines/spans), delete, and rewrite (rescue poorly written but informative pages).
  • Trained via supervised fine-tuning on targets curated from three teachers: Dripper (extraction), Qwen3.8-27B (refinement), and RePro (rewriting).
  • Pretraining 400M, 1.4B, and 2.8B models on ReScraper-curated data improves DCLM Core score by a relative 3.8–4.7% over the strongest baseline at each scale, including the multi-agent DataOrchestra pipeline.
  • Each operation contributes distinctively; unified extraction+cleaning outperforms cascaded separate models.

Introduction and Theoretical Foundation

LLM pretraining corpora rely heavily on web crawls, but raw HTML pages must be converted into clean, pretraining-ready text. This traditionally involves:

  1. Heuristic scrapers (e.g., resiliparse, trafilatura, jusText) that extract main content using hand-written rules over DOM structure.
  2. Rule-based filters (e.g., C4, RefinedWeb, FineWeb) that apply fixed thresholds on length, symbol ratios, and repetition.

Key limitations of the heuristic stack:

  • Rules mostly keep or drop whole pages, so a page with one noisy paragraph either loses useful content or retains noise.
  • Rules are inaccurate: 31–65% of pages dropped by each rule are judged worth keeping by an LLM-as-a-judge (Figure 1a).
  • Errors compound across stages—a scraper's mistakes cannot be undone by later filters.

Motivation for a unified model: While learned models have replaced individual stages (model-based scrapers, quality filters, editors, rewriters), each still consumes the output of the stage before it, preserving the cascade's weaknesses. The authors ask: Can a single small language model replace the entire pipeline, adaptively turning raw data into pretraining-ready text?


Methodology

Problem Formulation

ReScraper takes raw rendered HTML text (one block per line, each prefixed with <lid:N>) and generates a sequence of the form:

<extract> [extraction payload] (<keep> | <delete> | <edit> | <rewrite>) [operation payload]

The four operations are summarized below:

OperationPayloadExplanationTeacher
<extract>rm a-b, rm aAlways performed first. Drops lines that are not main content.Dripper
<keep>noneReturn extracted text unchanged.Qwen3.8-27B
<delete>noneDiscard page from corpus.Qwen3.8-27B
<edit>rm a-b, rm a, sub a: "s"Drop noise lines/spans; delete string s from line a.Qwen3.8-27B
<rewrite>rewritten documentReplace page with newly written text.RePro

Every operation except <rewrite> only removes text, keeping the output efficient (length grows with removals, not page length).

SFT Data Construction

Three teachers run sequentially on each page:

  1. Extraction teacher (Dripper): Identifies main content lines; removals form the <extract> payload.
  2. Refining teacher (Qwen3.8-27B): Judged with a fixed set of refinement rules. Pages become <delete> if valueless, <keep> if clean, or <edit> if noise lines/fragments are removed.
  3. Rewriting teacher (RePro): Pages scored ≥1.0 on FineWeb-Edu are rescued and rewritten instead of deleted, becoming <rewrite> targets.

Two-Stage Training

  • Stage 1: Teaches all operations, focusing on removal-based operations; <rewrite> is subsampled (0.1% of targets).
  • Stage 2: Strengthens <rewrite> learning with a mixture where <rewrite> comprises 30% of targets, encouraging the model to rewrite borderline pages.

Implementation details: Backbone is Qwen3-0.6B; context window 32,768 tokens; operation-tag loss weighted 5×; decoding at temperature 1.0, top-p 1.0.


Empirical Validation / Results

Main Results (Table 2)

Cleaning Method#Unique Tokens400M Core1B Core3B Core
Raw text17.69B0.133000.225440.28039
C4-rule5.06B0.137010.213620.24091
RefinedWeb-rule7.04B0.125720.253400.30058
FineWeb-rule5.16B0.146540.230220.27898
ProX-C11.80B0.142940.233910.29690
UltraX10.78B0.139460.261520.30642
DataOrchestra*13.60B0.146050.256480.30683
ReScraper (Ours)7.44B0.153450.273480.31841

*Multi-agent pipeline. Bold = best; underline = second-best.

Key findings:

  • ReScraper improves over the strongest rule-based baseline by relative 4.7% (400M), 7.9% (1B), 5.9% (3B).
  • Outperforms DataOrchestra (1.7B orchestrator + 0.6B/4B tools) by relative 5.1%, 6.6%, 3.8% at each scale, despite being a single 0.6B model.
  • Demonstrates weak-to-strong curation: a 0.6B model curates data that trains better 1.4B and 2.8B models.
  • ReScraper outperforms all scraper baselines (resiliparse, trafilatura, jusText, Dripper) with RefinedWeb-rule cleaning, and beats the Dripper+UltraX cascade by 6.6% (400M) and 3.2% (1B).

Ablations (Table 3, 1B scale)

VariantCore
Full ReScraper0.27348
w/o <edit>0.26696
w/o <rewrite>0.25221
w/o <delete>0.24951
w/o <extract>0.22806
Dripper <extract> (cascade)0.26363

Removing <extract> causes the largest drop (relative 16.6%); each operation contributes distinctively.

Quality and Diversity

  • ReScraper improves poor pages the most: raises DataMan by 1.28 on the poorest pages vs. at most 0.44 for baselines (Figure 6).
  • Keeps 95% of pages rated ≥1.0 by FineWeb-Edu, while RefinedWeb-rule drops 30% of them.
  • Retains the highest share of distinct 3-grams and 5-grams at equal token counts, confirming corpus diversity.

Learning from Teachers

  • Student follows teacher operations closely (token F1 of 89.3); operation distribution is similar.
  • Extraction matches Dripper (judge score 5.6 vs. 5.6).
  • Intentional deviation: student rewrites more borderline pages (FineWeb-Edu just below 1.0), retaining more tokens.

Theoretical and Practical Implications

Theoretical significance:

  • Demonstrates the feasibility of AI4AI (AI-for-AI) for pretraining data curation: a small learned model can replace an entire heuristic pipeline.
  • Shows that unified models outperform cascades: performing extraction and cleaning in one model beats running them as separate stages, because errors don't compound and the model can adapt to each page's content.
  • Provides evidence for weak-to-strong curation: curator size need not grow with the models it serves.

Practical implications:

  • ReScraper is more efficient than baselines: ~494 H200 GPU hours per pool vs. 1,366 for Dripper alone and 1,652 for DataOrchestra.
  • The model is open-sourced (model, dataset, and code are publicly released), enabling reproduction and further research.
  • The approach is scalable: a single 0.6B model can process large web corpora at a fraction of the cost of multi-model systems.

Conclusion

ReScraper introduces a unified, adaptive language model that replaces the heuristic scraping and cleaning stack for LLM pretraining data curation. Key takeaways:

  1. A single 0.6B model, trained on curated supervision from three teachers, produces higher-quality pretraining data than every baseline pipeline (rule-based, model-based, and multi-agent).
  2. The unified approach outperforms cascades of separate models, demonstrating that extraction and cleaning are best done together.
  3. Each operation (extract, keep, edit, delete, rewrite) plays a distinct and complementary role, with the model concentrating its changes on pages that need them most.
  4. The student faithfully follows its teachers at a fraction of their cost, with intentional deviations to improve borderline-page handling.

Future directions: The authors hope ReScraper motivates research toward curators that adapt data to the evolving needs of the model being trained, and ultimately toward recursive self-improvement, where each generation of models curates the pretraining data for the next.

Related papers