# From Search to Research: Exploring Search Scaling in Autonomous Quantitative Factor Mining

> Search scaling in autonomous research is capability-dependent: deeper search narrows cross-model gaps, but early trajectory state and parallel exploration critically shape final performance.

- **Source:** [arXiv](https://arxiv.org/abs/2609.35559)
- **Published:** 2026-10-03
- **Permalink:** https://picx.dev/p/P4N3xC
- **Whiteboard:** https://picx.dev/p/P4N3xC/image

## Summary

## Summary (Overview)

- **Novel Benchmark & Protocol**: Introduces a training-free, hypothesis-constrained protocol for measuring "search scaling" in autonomous research, using 50 quantitative factor-mining tasks grounded in financial research reports from the Chinese A-share market.
- **Key Finding on Capability vs. Search**: Initial performance is more strongly associated with base model capability, while deeper search can narrow cross-model gaps—weaker models can partially catch up with more iterations.
- **Model Grafting Insight**: Early research-state quality materially shapes final performance; transferring intermediate research states between models reveals that the initial trajectory direction is a critical determinant of outcome.
- **Parallel vs. Sequential**: Parallel search outperforms sequential search under the same iteration budget, consistent with the benefit of broader coverage of the search space.
- **Trajectory Analysis**: Higher-performing models more effectively diagnose failures, revise search directions, and preserve the intended economic hypothesis when selecting candidates, connecting qualitative behavior to quantitative outcomes.

---

## Introduction and Theoretical Foundation

### Background and Motivation

The paper extends the principle of **inference scaling**—where increased test-time computation improves LLM performance—to **autonomous LLM agents** through increased search budgets, termed **search scaling**. While prior work has characterized the mechanisms and limits of inference scaling in reasoning tasks, much less is known about these questions in the context of autonomous research, where agents must:

1. Observe experimental results
2. Diagnose failures
3. Choose the next experiment
4. Iteratively refine artifacts

The authors argue that search scaling is fundamentally different from simply generating more tokens or independent candidate answers: *"added computation pays off only when the agent converts feedback into useful experiments, so the return to search may depend on the base model."*

### Theoretical Foundation

The paper distinguishes three key quantities:

- **Base-model capability** $C_m$: inherent quality of the model
- **Nominal search budget** $b$: externally controlled number of research iterations
- **Realized resource consumption** $R_{m,i}(b)$: tokens, tool calls, elapsed time, and API cost incurred at depth $b$

The central question: *How does base-model capability shape the marginal value of additional experimentation and the performance level that iterative search can reach?*

### Why Quantitative Factor Mining?

The authors choose factor construction as a testbed because:

- It requires **end-to-end research loops** (hypothesis → implementation → evaluation → refinement)
- It is **training-free**—additional iterations change research behavior without increasing model-training compute
- **Economic hypotheses** constrain the search space, discouraging revisions that abandon the intended hypothesis
- **Noisy financial feedback** and local optima preserve a substantive generalization problem, making held-out evaluation essential

---

## Methodology

### Problem Formulation

Let $x_i$ denote an autonomous factor-research objective, $M_m$ a fixed base model, and $H$ a shared general-purpose agent harness. The system follows a continuous trajectory $\tau_{m,i}$ on task $i$. At search depth $b$, it selects an incumbent factor $f^{(b)}_{m,i}$ using development-period evidence. An independent evaluator measures held-out quality:

$$Q_{m,i}(b) = q(f^{(b)}_{m,i}, D_{\text{test}})$$

The checkpoint is an artifact produced by research, whereas $Q_{m,i}(b)$ is observed only by the evaluator.

### Controlled Iterative Search Protocol

- An **external harness loop** advances each model–task pair along a single research trajectory
- Each loop invocation permits one research cycle: act on current state, run experiments, inspect development-period feedback, update candidate
- The next cycle **inherits workspace artifacts, memory, and interaction history**
- The agent is **not told the eventual trajectory length**
- Checkpoints are frozen after every iteration; evaluation is separate from search and never feeds back into agent observations

### Measuring Search Scaling

Aggregate held-out outcomes over the same $N$ tasks at each depth:

$$\bar{Q}_m(b) = \frac{1}{N} \sum_{i=1}^{N} Q_{m,i}(b), \quad \Delta_m(b_1, b_2) = \bar{Q}_m(b_2) - \bar{Q}_m(b_1), \quad b_2 > b_1$$

The paper proposes a **shifted, saturating power law** to describe the budget-performance curve:

$$\bar{Q}_m(b) = Q_{m,\infty} - A_m (b + b_0)^{-\alpha_m}, \quad A_m > 0, \alpha_m > 0$$

Where:
- $Q_{m,\infty}$: plateau parameter
- $A_m$: scale of remaining gains
- $\alpha_m$: decay rate with depth
- $b_0$: shared offset fixed before fitting

### Experimental Setup

- **50 tasks**: grounded in quantitative research reports, spanning fundamental and technical ideas (momentum, reversal, valuation, earnings, growth, volatility, liquidity)
- **Environment**: Chinese A-share market
- **Development period**: Jan 1, 2013 – Jun 30, 2024
- **Evaluation period**: Jul 1, 2024 – Jul 17, 2026
- **Models**: 9 models evaluated along continuous 10-iteration trajectories
- **Interventions**: model grafting (transferring intermediate research states between models) and sequential–parallel comparison

---

## Empirical Validation / Results

### 1. Capability vs. Search Depth

The paper finds that:

- **Initial performance** is more strongly associated with model capability
- **Deeper search** can narrow cross-model gaps—weaker models benefit more from additional iterations in relative terms
- The marginal gains from search are not simply proportional to initial quality; the relationship between capability and search returns is nuanced

### 2. Model Grafting

Model grafting—transferring the research state (checkpoint) from one model to another mid-trajectory—reveals that:

- The **early research state materially shapes final performance**
- A strong model starting from a weak model's state does not fully recover the performance it would achieve from its own trajectory
- Conversely, a weak model starting from a strong model's state inherits significant benefits

This suggests that **path dependence** matters: the direction set early in the research process constrains later outcomes.

### 3. Parallel vs. Sequential Search

Under the same total iteration budget:

- **Parallel search** (running multiple independent trajectories and selecting the best) **outperforms sequential search** (one continuous trajectory)
- This is consistent with the benefit of **broader coverage of the search space**—parallel exploration reduces the risk of getting stuck in local optima

### 4. Trajectory Analysis

Higher-performing models exhibit qualitatively different behaviors:

- **More effective failure diagnosis**: they correctly attribute feedback to its causes
- **Better search redirection**: they revise search directions when evidence warrants
- **Hypothesis preservation**: they retain candidates consistent with the assigned economic hypothesis, avoiding "reward hacking" where revisions abandon the intended hypothesis to chase backtest scores

### Empirical Search Scaling Law

The fitted power law:

$$\bar{Q}_m(b) = Q_{m,\infty} - A_m (b + b_0)^{-\alpha_m}$$

shows **diminishing returns** with search depth across all models, but the rate of saturation ($\alpha_m$) and the plateau ($Q_{m,\infty}$) differ by model capability. Higher-capability models reach higher plateaus and exhibit more favorable cost-performance trade-offs.

---

## Theoretical and Practical Implications

### Theoretical Implications

1. **Search scaling is capability-dependent**: The marginal value of additional experimentation is not uniform across models. This extends the test-time scaling literature from independent sampling to dependent empirical search.

2. **Path dependence in research**: The model grafting results demonstrate that research trajectories are path-dependent—early decisions constrain later outcomes. This has implications for how we think about the "state" of an autonomous research agent.

3. **Coverage vs. depth trade-off**: The parallel-vs-sequential result provides evidence that search breadth can be as important as depth, aligning with theoretical results on exploration–exploitation trade-offs.

### Practical Implications

1. **Resource allocation**: Organizations deploying autonomous research agents should consider that weaker models may benefit more from additional search budget, while stronger models may warrant more parallel exploration.

2. **Adaptive policies**: The findings suggest that future progress requires *"stronger models together with adaptive policies for deploying test-time computation throughout the research process."*

3. **Evaluation design**: The protocol establishes a template for measuring search scaling that separates development-period search from held-out evaluation, avoiding feedback contamination.

4. **Risk of overfitting**: The hypothesis-constrained protocol highlights the risk of "adaptive selection that favors noise"—agents must be constrained to preserve intended economic hypotheses to avoid backtest overfitting.

---

## Conclusion

This paper establishes a rigorous framework for studying **search scaling** in autonomous research, using quantitative factor mining as a controlled yet realistic testbed. Key takeaways:

1. **Capability and search are complementary**: Initial quality tracks model capability, but deeper search can partially close cross-model gaps.

2. **Early state matters**: Model grafting shows that the research trajectory's early direction significantly shapes final outcomes, implying that initial planning quality is crucial.

3. **Breadth helps**: Parallel search outperforms sequential search under the same budget, favoring broader exploration.

4. **Behavioral differences drive outcomes**: Higher-performing models exhibit better failure diagnosis, search redirection, and hypothesis preservation.

### Future Directions

- **Adaptive budget allocation**: Developing policies that dynamically allocate search compute based on model capability and task difficulty
- **Better search organizations**: Exploring hybrid sequential–parallel strategies that combine the benefits of both
- **Stronger base models**: The ceiling on search scaling is partly set by base-model capability; improving base models remains essential
- **Generalization to other domains**: Extending the protocol to other autonomous research settings beyond finance

The paper concludes that *"future progress in autonomous research will require stronger models together with adaptive policies for deploying test-time computation throughout the research process."* The search scaling framework provides a foundation for understanding how these two levers interact.

---

_Markdown view of https://picx.dev/p/P4N3xC, served by PicX — AI-generated visual whiteboard summaries of research papers._
