# ReScraper: Unified Scraping and Cleaning of Web Data for Effective LLM Pretraining

> ReScraper, a single 0.6B language model, replaces the entire heuristic scraping and cleaning pipeline, improving LLM pretraining data quality by 3.8–4.7% over all baselines.

- **Source:** [arXiv](https://arxiv.org/abs/2609.34287)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/nq7Fk2
- **Whiteboard:** https://picx.dev/p/nq7Fk2/image

## Summary

## Summary (Overview)

- **ReScraper** is a unified 0.6B language model that replaces the entire heuristic pipeline of HTML scraping and rule-based cleaning for LLM pretraining data curation.
- It performs four operations in a single pass: **extract** (main content), **keep**, **edit** (remove noisy lines/spans), **delete**, and **rewrite** (rescue poorly written but informative pages).
- Trained via supervised fine-tuning on targets curated from three teachers: Dripper (extraction), Qwen3.8-27B (refinement), and RePro (rewriting).
- Pretraining 400M, 1.4B, and 2.8B models on ReScraper-curated data improves DCLM Core score by a **relative 3.8–4.7%** over the strongest baseline at each scale, including the multi-agent DataOrchestra pipeline.
- Each operation contributes distinctively; unified extraction+cleaning outperforms cascaded separate models.

---

## Introduction and Theoretical Foundation

LLM pretraining corpora rely heavily on web crawls, but raw HTML pages must be converted into clean, pretraining-ready text. This traditionally involves:

1. **Heuristic scrapers** (e.g., resiliparse, trafilatura, jusText) that extract main content using hand-written rules over DOM structure.
2. **Rule-based filters** (e.g., C4, RefinedWeb, FineWeb) that apply fixed thresholds on length, symbol ratios, and repetition.

**Key limitations of the heuristic stack:**
- Rules mostly keep or drop whole pages, so a page with one noisy paragraph either loses useful content or retains noise.
- Rules are inaccurate: 31–65% of pages dropped by each rule are judged worth keeping by an LLM-as-a-judge (Figure 1a).
- Errors compound across stages—a scraper's mistakes cannot be undone by later filters.

**Motivation for a unified model:** While learned models have replaced individual stages (model-based scrapers, quality filters, editors, rewriters), each still consumes the output of the stage before it, preserving the cascade's weaknesses. The authors ask: *Can a single small language model replace the entire pipeline, adaptively turning raw data into pretraining-ready text?*

---

## Methodology

### Problem Formulation

ReScraper takes raw rendered HTML text (one block per line, each prefixed with `<lid:N>`) and generates a sequence of the form:

```
<extract> [extraction payload] (<keep> | <delete> | <edit> | <rewrite>) [operation payload]
```

The four operations are summarized below:

| Operation | Payload | Explanation | Teacher |
|---|---|---|---|
| `<extract>` | `rm a-b`, `rm a` | Always performed first. Drops lines that are not main content. | Dripper |
| `<keep>` | none | Return extracted text unchanged. | Qwen3.8-27B |
| `<delete>` | none | Discard page from corpus. | Qwen3.8-27B |
| `<edit>` | `rm a-b`, `rm a`, `sub a: "s"` | Drop noise lines/spans; delete string `s` from line `a`. | Qwen3.8-27B |
| `<rewrite>` | rewritten document | Replace page with newly written text. | RePro |

Every operation except `<rewrite>` only removes text, keeping the output efficient (length grows with removals, not page length).

### SFT Data Construction

Three teachers run sequentially on each page:

1. **Extraction teacher (Dripper):** Identifies main content lines; removals form the `<extract>` payload.
2. **Refining teacher (Qwen3.8-27B):** Judged with a fixed set of refinement rules. Pages become `<delete>` if valueless, `<keep>` if clean, or `<edit>` if noise lines/fragments are removed.
3. **Rewriting teacher (RePro):** Pages scored ≥1.0 on FineWeb-Edu are rescued and rewritten instead of deleted, becoming `<rewrite>` targets.

### Two-Stage Training

- **Stage 1:** Teaches all operations, focusing on removal-based operations; `<rewrite>` is subsampled (0.1% of targets).
- **Stage 2:** Strengthens `<rewrite>` learning with a mixture where `<rewrite>` comprises 30% of targets, encouraging the model to rewrite borderline pages.

**Implementation details:** Backbone is Qwen3-0.6B; context window 32,768 tokens; operation-tag loss weighted 5×; decoding at temperature 1.0, top-p 1.0.

---

## Empirical Validation / Results

### Main Results (Table 2)

| Cleaning Method | #Unique Tokens | 400M Core | 1B Core | 3B Core |
|---|---|---|---|---|
| Raw text | 17.69B | 0.13300 | 0.22544 | 0.28039 |
| C4-rule | 5.06B | 0.13701 | 0.21362 | 0.24091 |
| RefinedWeb-rule | 7.04B | 0.12572 | 0.25340 | 0.30058 |
| FineWeb-rule | 5.16B | 0.14654 | 0.23022 | 0.27898 |
| ProX-C | 11.80B | 0.14294 | 0.23391 | 0.29690 |
| UltraX | 10.78B | 0.13946 | 0.26152 | 0.30642 |
| DataOrchestra* | 13.60B | 0.14605 | 0.25648 | 0.30683 |
| **ReScraper (Ours)** | **7.44B** | **0.15345** | **0.27348** | **0.31841** |

*Multi-agent pipeline. Bold = best; underline = second-best.

**Key findings:**
- ReScraper improves over the strongest rule-based baseline by **relative 4.7% (400M), 7.9% (1B), 5.9% (3B)**.
- Outperforms DataOrchestra (1.7B orchestrator + 0.6B/4B tools) by **relative 5.1%, 6.6%, 3.8%** at each scale, despite being a single 0.6B model.
- Demonstrates **weak-to-strong curation**: a 0.6B model curates data that trains better 1.4B and 2.8B models.
- ReScraper outperforms all scraper baselines (resiliparse, trafilatura, jusText, Dripper) with RefinedWeb-rule cleaning, and beats the Dripper+UltraX cascade by 6.6% (400M) and 3.2% (1B).

### Ablations (Table 3, 1B scale)

| Variant | Core |
|---|---|
| Full ReScraper | **0.27348** |
| w/o `<edit>` | 0.26696 |
| w/o `<rewrite>` | 0.25221 |
| w/o `<delete>` | 0.24951 |
| w/o `<extract>` | 0.22806 |
| Dripper `<extract>` (cascade) | 0.26363 |

Removing `<extract>` causes the largest drop (relative 16.6%); each operation contributes distinctively.

### Quality and Diversity

- ReScraper improves poor pages the most: raises DataMan by 1.28 on the poorest pages vs. at most 0.44 for baselines (Figure 6).
- Keeps 95% of pages rated ≥1.0 by FineWeb-Edu, while RefinedWeb-rule drops 30% of them.
- Retains the highest share of distinct 3-grams and 5-grams at equal token counts, confirming corpus diversity.

### Learning from Teachers

- Student follows teacher operations closely (token F1 of 89.3); operation distribution is similar.
- Extraction matches Dripper (judge score 5.6 vs. 5.6).
- Intentional deviation: student rewrites more borderline pages (FineWeb-Edu just below 1.0), retaining more tokens.

---

## Theoretical and Practical Implications

**Theoretical significance:**
- Demonstrates the feasibility of **AI4AI** (AI-for-AI) for pretraining data curation: a small learned model can replace an entire heuristic pipeline.
- Shows that **unified models outperform cascades**: performing extraction and cleaning in one model beats running them as separate stages, because errors don't compound and the model can adapt to each page's content.
- Provides evidence for **weak-to-strong curation**: curator size need not grow with the models it serves.

**Practical implications:**
- ReScraper is more efficient than baselines: ~494 H200 GPU hours per pool vs. 1,366 for Dripper alone and 1,652 for DataOrchestra.
- The model is open-sourced (model, dataset, and code are publicly released), enabling reproduction and further research.
- The approach is scalable: a single 0.6B model can process large web corpora at a fraction of the cost of multi-model systems.

---

## Conclusion

ReScraper introduces a unified, adaptive language model that replaces the heuristic scraping and cleaning stack for LLM pretraining data curation. Key takeaways:

1. A single 0.6B model, trained on curated supervision from three teachers, produces higher-quality pretraining data than every baseline pipeline (rule-based, model-based, and multi-agent).
2. The unified approach outperforms cascades of separate models, demonstrating that extraction and cleaning are best done together.
3. Each operation (extract, keep, edit, delete, rewrite) plays a distinct and complementary role, with the model concentrating its changes on pages that need them most.
4. The student faithfully follows its teachers at a fraction of their cost, with intentional deviations to improve borderline-page handling.

**Future directions:** The authors hope ReScraper motivates research toward curators that adapt data to the evolving needs of the model being trained, and ultimately toward **recursive self-improvement**, where each generation of models curates the pretraining data for the next.

---

_Markdown view of https://picx.dev/p/nq7Fk2, served by PicX — AI-generated visual whiteboard summaries of research papers._
