Full text not available for this paper
Summary (Overview)
- ReScraper is a unified 0.6B language model that replaces the entire heuristic pipeline of HTML scraping and rule-based cleaning for LLM pretraining data curation.
- It performs four operations in a single pass: extract (main content), keep, edit (remove noisy lines/spans), delete, and rewrite (rescue poorly written but informative pages).
- Trained via supervised fine-tuning on targets curated from three teachers: Dripper (extraction), Qwen3.8-27B (refinement), and RePro (rewriting).
- Pretraining 400M, 1.4B, and 2.8B models on ReScraper-curated data improves DCLM Core score by a relative 3.8–4.7% over the strongest baseline at each scale, including the multi-agent DataOrchestra pipeline.
- Each operation contributes distinctively; unified extraction+cleaning outperforms cascaded separate models.
Introduction and Theoretical Foundation
LLM pretraining corpora rely heavily on web crawls, but raw HTML pages must be converted into clean, pretraining-ready text. This traditionally involves:
- Heuristic scrapers (e.g., resiliparse, trafilatura, jusText) that extract main content using hand-written rules over DOM structure.
- Rule-based filters (e.g., C4, RefinedWeb, FineWeb) that apply fixed thresholds on length, symbol ratios, and repetition.
Key limitations of the heuristic stack:
- Rules mostly keep or drop whole pages, so a page with one noisy paragraph either loses useful content or retains noise.
- Rules are inaccurate: 31–65% of pages dropped by each rule are judged worth keeping by an LLM-as-a-judge (Figure 1a).
- Errors compound across stages—a scraper's mistakes cannot be undone by later filters.
Motivation for a unified model: While learned models have replaced individual stages (model-based scrapers, quality filters, editors, rewriters), each still consumes the output of the stage before it, preserving the cascade's weaknesses. The authors ask: Can a single small language model replace the entire pipeline, adaptively turning raw data into pretraining-ready text?
Methodology
Problem Formulation
ReScraper takes raw rendered HTML text (one block per line, each prefixed with <lid:N>) and generates a sequence of the form:
<extract> [extraction payload] (<keep> | <delete> | <edit> | <rewrite>) [operation payload]
The four operations are summarized below:
| Operation | Payload | Explanation | Teacher |
|---|---|---|---|
<extract> | rm a-b, rm a | Always performed first. Drops lines that are not main content. | Dripper |
<keep> | none | Return extracted text unchanged. | Qwen3.8-27B |
<delete> | none | Discard page from corpus. | Qwen3.8-27B |
<edit> | rm a-b, rm a, sub a: "s" | Drop noise lines/spans; delete string s from line a. | Qwen3.8-27B |
<rewrite> | rewritten document | Replace page with newly written text. | RePro |
Every operation except <rewrite> only removes text, keeping the output efficient (length grows with removals, not page length).
SFT Data Construction
Three teachers run sequentially on each page:
- Extraction teacher (Dripper): Identifies main content lines; removals form the
<extract>payload. - Refining teacher (Qwen3.8-27B): Judged with a fixed set of refinement rules. Pages become
<delete>if valueless,<keep>if clean, or<edit>if noise lines/fragments are removed. - Rewriting teacher (RePro): Pages scored ≥1.0 on FineWeb-Edu are rescued and rewritten instead of deleted, becoming
<rewrite>targets.
Two-Stage Training
- Stage 1: Teaches all operations, focusing on removal-based operations;
<rewrite>is subsampled (0.1% of targets). - Stage 2: Strengthens
<rewrite>learning with a mixture where<rewrite>comprises 30% of targets, encouraging the model to rewrite borderline pages.
Implementation details: Backbone is Qwen3-0.6B; context window 32,768 tokens; operation-tag loss weighted 5×; decoding at temperature 1.0, top-p 1.0.
Empirical Validation / Results
Main Results (Table 2)
| Cleaning Method | #Unique Tokens | 400M Core | 1B Core | 3B Core |
|---|---|---|---|---|
| Raw text | 17.69B | 0.13300 | 0.22544 | 0.28039 |
| C4-rule | 5.06B | 0.13701 | 0.21362 | 0.24091 |
| RefinedWeb-rule | 7.04B | 0.12572 | 0.25340 | 0.30058 |
| FineWeb-rule | 5.16B | 0.14654 | 0.23022 | 0.27898 |
| ProX-C | 11.80B | 0.14294 | 0.23391 | 0.29690 |
| UltraX | 10.78B | 0.13946 | 0.26152 | 0.30642 |
| DataOrchestra* | 13.60B | 0.14605 | 0.25648 | 0.30683 |
| ReScraper (Ours) | 7.44B | 0.15345 | 0.27348 | 0.31841 |
*Multi-agent pipeline. Bold = best; underline = second-best.
Key findings:
- ReScraper improves over the strongest rule-based baseline by relative 4.7% (400M), 7.9% (1B), 5.9% (3B).
- Outperforms DataOrchestra (1.7B orchestrator + 0.6B/4B tools) by relative 5.1%, 6.6%, 3.8% at each scale, despite being a single 0.6B model.
- Demonstrates weak-to-strong curation: a 0.6B model curates data that trains better 1.4B and 2.8B models.
- ReScraper outperforms all scraper baselines (resiliparse, trafilatura, jusText, Dripper) with RefinedWeb-rule cleaning, and beats the Dripper+UltraX cascade by 6.6% (400M) and 3.2% (1B).
Ablations (Table 3, 1B scale)
| Variant | Core |
|---|---|
| Full ReScraper | 0.27348 |
w/o <edit> | 0.26696 |
w/o <rewrite> | 0.25221 |
w/o <delete> | 0.24951 |
w/o <extract> | 0.22806 |
Dripper <extract> (cascade) | 0.26363 |
Removing <extract> causes the largest drop (relative 16.6%); each operation contributes distinctively.
Quality and Diversity
- ReScraper improves poor pages the most: raises DataMan by 1.28 on the poorest pages vs. at most 0.44 for baselines (Figure 6).
- Keeps 95% of pages rated ≥1.0 by FineWeb-Edu, while RefinedWeb-rule drops 30% of them.
- Retains the highest share of distinct 3-grams and 5-grams at equal token counts, confirming corpus diversity.
Learning from Teachers
- Student follows teacher operations closely (token F1 of 89.3); operation distribution is similar.
- Extraction matches Dripper (judge score 5.6 vs. 5.6).
- Intentional deviation: student rewrites more borderline pages (FineWeb-Edu just below 1.0), retaining more tokens.
Theoretical and Practical Implications
Theoretical significance:
- Demonstrates the feasibility of AI4AI (AI-for-AI) for pretraining data curation: a small learned model can replace an entire heuristic pipeline.
- Shows that unified models outperform cascades: performing extraction and cleaning in one model beats running them as separate stages, because errors don't compound and the model can adapt to each page's content.
- Provides evidence for weak-to-strong curation: curator size need not grow with the models it serves.
Practical implications:
- ReScraper is more efficient than baselines: ~494 H200 GPU hours per pool vs. 1,366 for Dripper alone and 1,652 for DataOrchestra.
- The model is open-sourced (model, dataset, and code are publicly released), enabling reproduction and further research.
- The approach is scalable: a single 0.6B model can process large web corpora at a fraction of the cost of multi-model systems.
Conclusion
ReScraper introduces a unified, adaptive language model that replaces the heuristic scraping and cleaning stack for LLM pretraining data curation. Key takeaways:
- A single 0.6B model, trained on curated supervision from three teachers, produces higher-quality pretraining data than every baseline pipeline (rule-based, model-based, and multi-agent).
- The unified approach outperforms cascades of separate models, demonstrating that extraction and cleaning are best done together.
- Each operation (extract, keep, edit, delete, rewrite) plays a distinct and complementary role, with the model concentrating its changes on pages that need them most.
- The student faithfully follows its teachers at a fraction of their cost, with intentional deviations to improve borderline-page handling.
Future directions: The authors hope ReScraper motivates research toward curators that adapt data to the evolving needs of the model being trained, and ultimately toward recursive self-improvement, where each generation of models curates the pretraining data for the next.
Related papers
- Shortcutting the Fix: Agentic Shortcutting in Software-Engineering Benchmarks
Agentic shortcutting—agents exploiting leaked solutions like upstream repos or Git history—inflates SWE benchmark scores by up to 82%, but a simple originality prompt cuts exploitation to under 11%.
- ScAn-Bench: Evaluating Scaling Analysis Methodology
ScAn-Bench introduces the first surrogate benchmarks for scaling analysis, revealing no universally optimal data acquisition and extrapolation strategy exists across LLMs and VLMs.
- How Far Are We from Removing the Visual Encoder? Scaling Laws for Encoder-Free Multimodal Pretraining
Encoder-free multimodal LLMs match encoder-based performance at ~10^22 FLOPs, shifting compute-optimal allocation toward larger decoders and enabling viable encoder-free pretraining.